Lightning AI

Infrastructure Operations Engineer (APAC)

Lightning AI · Singapore

Posted about 2 months ago

or apply directly on Lightning AI's site. We never take the application ourselves.

Is this posting real?

This role has been open
66 days
Lightning AI's roles stay open a median of 66 days
Reposted
No
Salary listed
No
6% of Lightning AI's roles list one
Ghost-job risk at Lightning AI
high
47 stale, 1 reposted of 54 open
Hiring momentum
61 roles opened in the last 90 days
↑ up vs. the prior 90 days
Last confirmed on the employer's board
2026-10-08

Measured from postings appearing on and disappearing from Lightning AI's own greenhouse board since 2026-08-03. Full hiring picture for Lightning AI.

About this role

As a Senior Infrastructure Operations Engineer at Lightning AI, you will be responsible for operating and troubleshooting large-scale GPU and bare-metal infrastructure. Your role involves resolving complex infrastructure issues, building automation tools, and improving operational processes while collaborating with various engineering teams. You will also participate in an on-call rotation to support production infrastructure.

Our read on this posting3.2out of 5
benefits
5/5
freshness
1/5
career value
5/5
role clarity
5/5
pay transparency
0/5

Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Lightning AI as an employer.

What you need

  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems.
  • Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.
  • Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems.
  • Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure.
  • Strong systems and networking fundamentals, with a track record of owning production issues through resolution.

Nice to have

  • Experience troubleshooting, provisioning, or operating bare-metal server infrastructure.
  • Experience operating or troubleshooting GPU, HPC, or other high-performance compute infrastructure.
  • Familiarity with NVIDIA GPUs, DCGM, InfiniBand, RoCE/RDMA, NVLink, or high-speed data center networking.
  • Experience with hardware management and provisioning technologies such as PXE, BMC, IPMI, Redfish, or iDRAC.
  • Experience with distributed or high-performance storage systems such as VAST, Ceph, GPFS, or WEKA.

What you get

  • Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
  • Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
  • Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
  • Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
  • Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
  • Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.

Worth weighing

  • Role requires participation in an on-call rotation, which may affect work-life balance.
  • Position is remote but candidates must reside in Singapore, limiting geographic flexibility.
  • The job involves troubleshooting and resolving complex infrastructure issues, which may lead to high-pressure situations.

Summarised from Lightning AI's posting. Read the full original.

Listed by Lightning AI on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.

More roles at Lightning AI

All 54 open roles at Lightning AI →