Infrastructure Engineer (GPU & Compute)
Lightning AI · London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
Posted about 2 months ago
or apply directly on Lightning AI's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 66 days Lightning AI's roles stay open a median of 66 days
- Reposted
- No
- Salary listed
- No 6% of Lightning AI's roles list one
- Ghost-job risk at Lightning AI
- high 47 stale, 1 reposted of 54 open
- Hiring momentum
- 61 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-10-08
Measured from postings appearing on and disappearing from Lightning AI's own greenhouse board since 2026-08-03. Full hiring picture for Lightning AI.
About this role
As a Senior GPU & Compute Infrastructure Engineer at Lightning AI, you will be responsible for the validation and operation of large-scale bare-metal compute infrastructure, focusing on GPU-enabled systems. Your day-to-day tasks will include managing system diagnostics, developing automation tools, and ensuring that clusters are prepared for demanding AI/ML and HPC workloads. You will collaborate with various teams to enhance infrastructure reliability and performance.
- benefits
- 5/5
- freshness
- 1/5
- career value
- 4/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Lightning AI as an employer.
What you need
- 5+ years of experience in infrastructure engineering, systems engineering, or related roles
- Strong Linux systems experience in production environments
- Hands-on experience with GPU-enabled systems and tools such as NVIDIA DCGM
- Familiarity with bare-metal provisioning and system bring-up workflows
- Proficiency in Python or similar scripting/programming languages for automation
- Ability to debug complex issues across hardware, OS, GPUs, and system software
Nice to have
- Experience with high-performance interconnects (e.g., InfiniBand, NVLink)
- Experience with PXE boot environments, LiveCD systems, or image-based provisioning workflows
- Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
- Data center operations experience, including working with physical hardware
- Experience supporting AI/ML or HPC workloads at scale
What you get
- Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
- Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
- Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
- Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
- Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
- Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.
Worth weighing
- No salary listed in the posting, but a range is provided.
- The role may be fully remote or hybrid, but visa sponsorship is not available.
- The position involves working with both hardware and software, which may require a broad skill set.
- The company emphasizes a fast-paced environment, which may not suit everyone.
Summarised from Lightning AI's posting. Read the full original.
Listed by Lightning AI on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.
More roles at Lightning AI
- Infrastructure Engineer (Storage)London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
- Infrastructure Operations Engineer (APAC)Singapore
- Infrastructure Operations EngineerLondon, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
- AI Platform Support Engineer (EMEA)London, England, United Kingdom
- Director of Customer ExperienceNew York, New York, United States
- Senior Network EngineerRemote
- Global Tax LeadSan Francisco, California, United States
- Technical People Operations SpecialistNew York, New York, United States