
Senior Software Engineer - NVLink Rack Scale Stability and Reliability
NVIDIA · Massachusetts, United States
- Hybrid
- Full-time
- $200,000 / year
- Massachusetts, United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Engineer NVLink Rack-Scale Systems stability.
- Develop diagnostics, automation, and infrastructure.
- Lead reliability validation and issue resolution.
- Triage complex multi-domain system issues.
- Collaborate across multiple engineering teams.
About the role
About NVIDIA
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.Job Summary
We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely with architects and developers building our next-generation NVLink and NVSwitch systems, helping transform first-of-their-kind platforms into stable, reliable, and volume production-ready systems. You will work on complex system-level challenges spanning resiliency, diagnostics, recovery, and large-scale AI infrastructure, contributing directly to the software foundation powering next-generation datacenter deployments.What You Will Be Doing
- Drive platform bringup, feature enablement, end-to-end software validation, and debug for next-generation NVLink-based GPU and rack-scale systems.
- Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.
- Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.
- Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.
- Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability.
- Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.
- Create automation, dashboards, runbooks, and debug workflows that improve root-cause analysis and operational efficiency.
What We Need To See
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
- 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.
- Strong programming skills in C/C++ and Python; Bash/Shell scripting experience is a plus.
- Strong system-level debugging across software, firmware, hardware, and networking layers.
- Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.
- Experience with large-scale AI systems, including platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.
- Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.
- Strong communication and collaboration skills across engineering, customer, and operations teams.
- Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.
Ways To Stand Out From The Crowd
- Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters such as NVIDIA GB200 NVL72.
- Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training/inference systems.
- Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.
- Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.
Compensation and Benefits
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits.Key skills/competency
- Senior Software Engineer
- NVLink Rack-Scale Systems
- Stability and Reliability
- System Software
- Firmware
- Networking
- Data Center Infrastructure
- Distributed Systems
- C/C++
- Python
- AI Infrastructure
- System-Level Debugging
Skills & topics
- Senior Software Engineer
- NVLink
- Rack Scale
- Stability
- Reliability
- System Software
- Firmware
- Networking
- Data Center
- AI Infrastructure
- C++
- Python
- Debugging
- NVIDIA
How to get hired
- Tailor your resume: Highlight experience with C/C++, Python, system-level debugging, and large-scale AI systems.
- Showcase relevant projects: Detail your contributions to platform bringup, reliability engineering, and automation.
- Prepare for technical interviews: Expect questions on distributed systems, networking, and debugging complex issues.
- Demonstrate collaboration skills: Emphasize your ability to work effectively with cross-functional teams.
Technical preparation
Master C/C++ and Python for system software.,Practice system-level debugging across layers.,Review networking protocols and fabric analysis.,Study large-scale AI system architectures.
Behavioral questions
Describe a complex system failure you resolved.,How do you handle conflicting priorities?,Tell me about a time you collaborated effectively.,How do you approach building reliable systems?
Frequently asked questions
- What specific NVIDIA technologies are most relevant for this Senior Software Engineer role?
- For this Senior Software Engineer position at NVIDIA, demonstrating experience with NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters like the NVIDIA GB200 NVL72 would be highly advantageous. Familiarity with NVIDIA GPU systems is also a significant plus.
- What is the expected level of system-level debugging expertise for this role?
- This Senior Software Engineer role requires strong system-level debugging skills across software, firmware, hardware, and networking layers. Candidates should be adept at triaging complex, multi-domain issues using logs, telemetry, and structured debugging methods.
- How does NVIDIA approach reliability and MTBI validation for its systems?
- NVIDIA approaches reliability and MTBI validation through rigorous stress testing, telemetry analysis, failure injection, and systematic issue resolution. This Senior Software Engineer role plays a key part in these validation efforts.
- What programming languages and scripting skills are essential for this Senior Software Engineer position?
- Strong programming skills in C/C++ and Python are essential for this role. While not strictly required, Bash/Shell scripting experience is considered a plus for developing automation and tools.
- What networking fundamentals are critical for the Senior Software Engineer role at NVIDIA?
- Solid networking fundamentals are critical, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis. This knowledge is vital for ensuring the stability and reliability of NVLink rack-scale systems.
- How can I best highlight my experience with large-scale AI systems for this role?
- To highlight your experience with large-scale AI systems, focus on your involvement in platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging. Quantify your impact where possible.
- What is NVIDIA's stance on using AI tools in the recruiting process for this Senior Software Engineer role?
- NVIDIA utilizes AI tools in its recruiting processes. Candidates should be aware that AI may be involved in various stages of evaluation for this Senior Software Engineer position.
- What is the application deadline for the Senior Software Engineer - NVLink Rack Scale Stability position?