Infrastructure Operations Engineer
Lightning AI · London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
Posted about 2 months ago
or apply directly on Lightning AI's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 66 days Lightning AI's roles stay open a median of 66 days
- Reposted
- No
- Salary listed
- No 6% of Lightning AI's roles list one
- Ghost-job risk at Lightning AI
- high 47 stale, 1 reposted of 54 open
- Hiring momentum
- 61 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-10-08
Measured from postings appearing on and disappearing from Lightning AI's own greenhouse board since 2026-08-03. Full hiring picture for Lightning AI.
About this role
As a Senior Infrastructure Operations Engineer at Lightning AI, you will operate and troubleshoot large-scale GPU and bare-metal infrastructure, focusing on Linux, networking, storage, and cluster orchestration. You will serve as a technical escalation point for complex issues, drive operational workflows, and build automation to enhance efficiency. The role emphasizes ownership of problems from investigation to resolution and collaboration across various engineering and operational teams.
- benefits
- 5/5
- freshness
- 1/5
- career value
- 4/5
- role clarity
- 5/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Lightning AI as an employer.
What you need
- Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
- Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems.
- Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.
- Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems.
- Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure.
- Strong systems and networking fundamentals, with a track record of owning production issues through resolution.
Nice to have
- Experience troubleshooting, provisioning, or operating bare-metal server infrastructure.
- Experience operating or troubleshooting GPU, HPC, or other high-performance compute infrastructure.
- Familiarity with NVIDIA GPUs, DCGM, InfiniBand, RoCE/RDMA, NVLink, or high-speed data center networking.
- Experience with hardware management and provisioning technologies such as PXE, BMC, IPMI, Redfish, or iDRAC.
- Experience with distributed or high-performance storage systems such as VAST, Ceph, GPFS, or WEKA.
What you get
- Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
- Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
- Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
- Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
- Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
- Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.
Worth weighing
- No salary listed in the posting, but a range is provided.
- The role may require participation in an on-call rotation, which could impact work-life balance.
- The position is remote-friendly, but candidates must be located within the U.S. and visa sponsorship is not available.
Summarised from Lightning AI's posting. Read the full original.
Listed by Lightning AI on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.
More roles at Lightning AI
- Infrastructure Engineer (GPU & Compute)London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
- Infrastructure Engineer (Storage)London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
- Infrastructure Operations Engineer (APAC)Singapore
- AI Platform Support Engineer (EMEA)London, England, United Kingdom
- Director of Customer ExperienceNew York, New York, United States
- Senior Network EngineerRemote
- Global Tax LeadSan Francisco, California, United States
- Technical People Operations SpecialistNew York, New York, United States