or apply directly on SpaceXAI's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 14 days SpaceXAI's roles stay open a median of 45 days
- Reposted
- No
- Salary listed
- No 32% of SpaceXAI's roles list one
- Ghost-job risk at SpaceXAI
- low 0 stale, 5 reposted of 258 open
- Hiring momentum
- 339 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-09-17
Measured from postings appearing on and disappearing from SpaceXAI's own greenhouse board since 2026-08-03. Full hiring picture for SpaceXAI.
About this role
As a Site Reliability Engineer at SpaceXAI, you will focus on ensuring the reliability of campus operations by designing monitoring systems, leading incident responses, and managing cross-functional reliability projects. Your role will involve technical leadership during incidents, maintaining playbooks, and collaborating with teams across various disciplines such as compute, network, and storage. You will also participate in on-call rotations and drive improvements based on incident feedback.
- benefits
- 1/5
- freshness
- 4/5
- career value
- 4/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of SpaceXAI as an employer.
What you need
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments
- Proven large-scale incident command experience and calm technical leadership on a bridge
- Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
- Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner
Nice to have
- Experience in AI/ML infrastructure or supercomputing environments
- Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries
- Experience running game days, dependency mapping, and closed-loop corrective action programs
- Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry
- Prior work in a fast-paced startup or tech company like SpaceXAI
Worth weighing
- No salary listed
- The role requires calm incident leadership and technical command during high-pressure situations
- The position involves a significant amount of collaboration across various technical disciplines
- Expectations for strong communication skills and the ability to share knowledge concisely and accurately with teammates
Summarised from SpaceXAI's posting. Read the full original.
Listed by SpaceXAI on their greenhouse job board, last confirmed open on 2026-09-17. PitchMeAI is not the employer.
More roles at SpaceXAI
- Lead, Driver - MemphisSouthaven, MS; Memphis, TN
- Supervisor, Logistics (Second Shift) - MemphisSouthaven, MS; Memphis, TN
- Program Manager, Prohibited & Regulated ContentPalo Alto, CA; Bastrop, TX; New York, NY
- Lead, Inventory Specialist - MemphisSouthaven, MS; Memphis, TN
- Program Manager, Harmful ActivityPalo Alto, CA; Bastrop, TX; New York, NY
- Electrician, Operations - MemphisMemphis, TN
- Lead, Receiving Specialist - MemphisSouthaven, MS; Memphis, TN
- Power System Engineer (Transmission & Distribution) - MemphisMemphis, TN