Lightning AI

Infrastructure Operations Engineer

Lightning AI · London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States

Posted about 2 months ago

or apply directly on Lightning AI's site. We never take the application ourselves.

Is this posting real?

This role has been open
66 days
Lightning AI's roles stay open a median of 66 days
Reposted
No
Salary listed
No
6% of Lightning AI's roles list one
Ghost-job risk at Lightning AI
high
47 stale, 1 reposted of 54 open
Hiring momentum
61 roles opened in the last 90 days
↑ up vs. the prior 90 days
Last confirmed on the employer's board
2026-10-08

Measured from postings appearing on and disappearing from Lightning AI's own greenhouse board since 2026-08-03. Full hiring picture for Lightning AI.

About this role

As a Senior Infrastructure Operations Engineer at Lightning AI, you will operate and troubleshoot large-scale GPU and bare-metal infrastructure, focusing on Linux, networking, storage, and cluster orchestration. You will serve as a technical escalation point for complex issues, drive operational workflows, and build automation to enhance efficiency. The role emphasizes ownership of problems from investigation to resolution and collaboration across various engineering and operational teams.

Our read on this posting3.0out of 5
benefits
5/5
freshness
1/5
career value
4/5
role clarity
5/5
pay transparency
0/5

Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Lightning AI as an employer.

What you need

  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems.
  • Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.
  • Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems.
  • Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure.
  • Strong systems and networking fundamentals, with a track record of owning production issues through resolution.

Nice to have

  • Experience troubleshooting, provisioning, or operating bare-metal server infrastructure.
  • Experience operating or troubleshooting GPU, HPC, or other high-performance compute infrastructure.
  • Familiarity with NVIDIA GPUs, DCGM, InfiniBand, RoCE/RDMA, NVLink, or high-speed data center networking.
  • Experience with hardware management and provisioning technologies such as PXE, BMC, IPMI, Redfish, or iDRAC.
  • Experience with distributed or high-performance storage systems such as VAST, Ceph, GPFS, or WEKA.

What you get

  • Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
  • Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
  • Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
  • Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
  • Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
  • Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.

Worth weighing

  • No salary listed in the posting, but a range is provided.
  • The role may require participation in an on-call rotation, which could impact work-life balance.
  • The position is remote-friendly, but candidates must be located within the U.S. and visa sponsorship is not available.

Summarised from Lightning AI's posting. Read the full original.

Listed by Lightning AI on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.

More roles at Lightning AI

All 54 open roles at Lightning AI →