
Staff Cloud Infrastructure Engineer (up to $200k)
Dex · United States
- Hybrid
- Full-time
- $200,000 / year
- United States
Job highlights
- Lead cloud infrastructure design for SaaS and customer-managed deployments.
- Architect scalable, reliable multi-tenant systems.
- Define and implement observability and automation strategies.
- Hands-on role from design to production operations.
- Opportunity to shape foundational systems in early-stage company.
About the role
About the Role
This company is building a lightweight runtime that turns AI agents, workflows, and backend services into durable processes. Engineers can focus on business logic, not failure mechanics. The runtime ships as a single Rust binary with a custom storage layer and low-latency orchestration, already running critical financial workflows for Fortune 500 enterprises and Tier 1 banks. Founded by creators of Apache Flink and leaders from Meta's messaging infrastructure, it's backed by leading VCs and angels.
As a Staff Cloud Infrastructure Engineer, you'll take staff-level ownership across the entire cloud infrastructure surface. This means setting the design direction for a managed multi-tenant SaaS offering, as well as for the infrastructure running inside customer cloud accounts for BYOC and on-prem installations. You'll define fleet-wide reliability and observability, building the automation to scale across deployment methods. This isn't just operating existing infrastructure; it's architecting and building the foundational systems for a product that spans multiple deployment models, with significant design authority in a small, high-impact team.
The Work
- Design and implement infrastructure for a multi-tenant SaaS offering, covering control plane, networking, storage, and observability.
- Architect and build infrastructure for customer-managed deployments (BYOC, on-prem), ensuring consistency and scalability.
- Define and implement fleet-wide reliability and observability standards, including SLOs, metrics, tracing, logging, and runbooks.
- Develop automation to scale infrastructure across diverse deployment methods and cloud providers.
- Participate in the cloud on-call rotation, providing critical timezone coverage for a global user base.
What You Bring
- Operated production SaaS or platform infrastructure at staff level, with strong opinions on multi-tenant systems formed from real failure modes.
- Deep working knowledge of at least one major cloud provider's architecture; experience with multiple is a plus.
- Designed and run large-scale Kubernetes-based stateful workloads, balancing continuous delivery with safety.
- Fluency in a systems language (Rust, Go, or C++); willingness to learn Rust on the job.
- Hands-on at staff level: comfortable taking a problem from architectural design through to production operations.
- US-based, able to join a cloud on-call rotation for timezone coverage.
Key skills/competency
- Cloud Infrastructure
- SaaS Offering
- Kubernetes
- Reliability
- Observability
- Automation
- Multi-tenant Systems
- Systems Language (Rust, Go, C++)
- Cloud Architecture
- Staff-Level Ownership
Skills & topics
- Staff Cloud Infrastructure Engineer
- Cloud Infrastructure
- SaaS
- Kubernetes
- Reliability
- Observability
- Automation
- Rust
- Go
- C++
- Multi-tenant Systems
- BYOC
- On-prem
- Systems Engineer
- Platform Engineer
How to get hired
- Tailor your resume: Highlight your staff-level experience in production SaaS infrastructure and multi-tenant systems.
- Showcase cloud expertise: Emphasize your deep knowledge of major cloud providers and Kubernetes.
- Demonstrate systems language fluency: Detail your experience with Rust, Go, or C++, and willingness to learn.
- Prepare for technical discussions: Be ready to discuss architectural designs and operational challenges.
- Apply through Dex: Utilize the Dex platform for a streamlined application process.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the primary technology stack for this Staff Cloud Infrastructure Engineer role?
- The primary technology stack involves cloud infrastructure (AWS, Azure, or GCP), Kubernetes for large-scale stateful workloads, and systems programming languages like Rust, Go, or C++. The company is particularly interested in candidates willing to learn Rust. You'll also be working with custom storage layers and low-latency orchestration.
- What kind of team will I be joining as a Staff Cloud Infrastructure Engineer at this company?
- You will join a small, high-impact team founded by experienced individuals, including creators of Apache Flink and leaders from Meta's messaging infrastructure. The role offers significant design authority and the opportunity to shape foundational systems.
- What are the key responsibilities for a Staff Cloud Infrastructure Engineer at this company?
- Key responsibilities include designing and implementing infrastructure for a multi-tenant SaaS offering, architecting customer-managed deployments (BYOC, on-prem), defining fleet-wide reliability and observability, and developing automation for scaling infrastructure across various cloud providers and deployment methods.
- What distinguishes this Staff Cloud Infrastructure Engineer role from others?
- This role is a rare, high-impact staff position at an early-stage company solving a complex problem with a novel runtime. You'll have significant design authority and the opportunity to build foundational systems for a product spanning multiple deployment models.
- How does Dex facilitate the application process for the Staff Cloud Infrastructure Engineer position?
- Dex acts as an AI recruiter agent. By signing up on their platform, you can provide your technical stack and preferences. Dex will manage your applications, provide a brief on the company and role, and streamline the process, helping you apply directly and get matched with similar opportunities.
- Is this Staff Cloud Infrastructure Engineer role remote or on-site?
- The job description specifies that candidates must be US-based and able to join a cloud on-call rotation for timezone coverage. While not explicitly stated as remote, the emphasis on US-based and timezone coverage suggests a strong possibility of remote work flexibility within the US.