PitchMeAI
Babylist

Staff Site Reliability Engineer

Babylist · United States

  • Hybrid
  • Full-time
  • $249,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Own AWS infrastructure and reliability for 9M+ users.
  • Evolve CI/CD, AWS, and developer tooling.
  • Support developers across all environments.
  • Lead incident response and platform strategy.
  • Shape infrastructure for a growing tech company.

About the role

About Babylist

Babylist is the leading platform for expecting and new families. More than 10 million people shop with Babylist every year, making it the go-to destination for seamless purchasing, guidance, and expert recommendations. As a modern, AI-forward tech company, Babylist has expanded from a universal registry into a full ecosystem — the Babylist Shop, Babylist Health, Babylist Money, NYC and LA showrooms, branded content, and more — generating $750M in revenue in 2025. Building the generational brand in baby, Babylist is reshaping the $235B kids and baby market and helping parents feel confident, connected, and cared for at every step.

Our Ways of Working

Babylist is remote-first with team members across the U.S. and Canada who move fast, think smart, and use AI as part of how they work every day — not as an experiment, as an expectation. We come together twice a year to build the relationships behind the work, and we hire people who are genuinely excited about what's possible and prove it through how they show up.

How We Build

Babylist is in the middle of a fundamental shift in how software gets made, and we are not tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI changes everything. How teams are structured, how decisions get made, how fast ideas become working software. Our engineers own problems end to end, working directly with product, design, and business partners with short feedback loops and real stakeholder access. We ship, learn, and iterate fast. When something is not working, we throw it out and start over — project failure and personal failure are not the same thing here. AI tools are as natural to our workflow as an IDE or version control. We are not exploring this, we are living it. Our engineers use AI to explore tradeoffs, pressure-test designs, and move from problem to solution in hours instead of days. They generate code with AI so they can stay focused on the decisions that actually require human judgment — not the routine ones. More velocity means more time for craft: better test coverage, stronger architecture, and deeper customer understanding. We hold ourselves to a higher quality bar because of AI, not in spite of it. We are building this playbook in real time, and we are looking for people who want to build it with us. If you have already changed how you work because of AI — or you are ready to — and you care more about shipping something great than following a prescribed process, we should talk.

Our Tech Stack

  • Ruby on Rails
  • AWS
  • Sidekiq
  • MySQL
  • Redis

What the Role Is

Babylist's Platform team is the foundation every engineering team builds on — and this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'll own the infrastructure and reliability practices that support 9 million+ users and the engineers who build for them. Babylist started as an e-commerce and registry platform, and we're actively growing beyond that — into health, media, mobile, and new product surfaces that don't exist yet. The Platform team is the foundation that makes all of it possible. This isn't a maintenance role — you'll be actively evolving how we build and operate AWS infrastructure, CI systems, and developer tooling. You'll work cross-functionally across all of Babylist Engineering, which means your decisions have wide leverage.

Who You Are

  • Deep hands-on Terraform expertise — you own IaC, not just contribute to it
  • Proven AWS experience at scale — EKS, RDS, cloud networking, DNS, CDNs, load balancers — you know the gotchas
  • Experienced operating Kubernetes in production — you've debugged the hard stuff, not just deployed the easy stuff
  • Comfortable designing and improving CI/CD systems — CircleCI, GitHub Actions, or similar; you care about developer velocity, not just pipeline uptime
  • Strong observability instincts — Datadog, Sentry, PagerDuty, Cronitor — you build alerting that's actionable, not noisy
  • Experienced with on-call and incident management — you've run the post-mortems and actually changed things afterward
  • Comfortable supporting developers across local, staging, and production — you're a resource, not a gatekeeper
  • You naturally reach for AI in your work — at Babylist, every team uses AI daily. You're already using it to move faster and improve your output, and you stay curious about what's coming next.

How You Will Make An Impact

  • Infrastructure ownership — manage and evolve our AWS environment using Terraform, keeping EKS clusters, databases, and core services current and performant
  • CI/CD reliability — own the speed and reliability of our CI systems for the full Engineering org — every deploy starts here
  • Developer support — be the person engineers turn to when environments break; unblock them fast across local, staging, and production
  • Monitoring & alerting standards — establish and socialize best practices so the right people get paged for the right reasons
  • Incident response — lead or support incident response, drive post-incident reviews, and close the loop so the same thing doesn't happen twice
  • Platform strategy — contribute to architectural decisions that shape how Babylist's infrastructure evolves over the next several years

Why This Role

  • Platform is the team every engineering team depends on — your work has outsized leverage across the entire product org, not just one area
  • The infrastructure is solid but actively evolving — you're not inheriting chaos, you're shaping what comes next
  • This is a staff-level role with real cross-team visibility — you'll influence how Babylist engineers build and ship, not just keep the lights on
  • You'll work on systems that support millions of families at a high-stakes life moment — the scale is real and the product context makes the reliability work matter

About Compensation

We use a market-based approach to compensation. The starting salary range for this role is:$226,673 to $271,991Your starting salary will be based on your location, experience, and qualifications, with increases over time tied to performance, role growth, and internal pay equity.

Key skills/competency

  • Site Reliability Engineering (SRE)
  • Infrastructure as Code (IaC)
  • Terraform
  • Amazon Web Services (AWS)
  • Kubernetes (EKS)
  • CI/CD
  • Observability
  • Incident Management
  • Developer Tooling
  • Artificial Intelligence (AI)

Skills & topics

  • Site Reliability Engineer
  • SRE
  • Staff Engineer
  • Infrastructure
  • AWS
  • Kubernetes
  • Terraform
  • CI/CD
  • Observability
  • Incident Management
  • Remote

How to get hired

  • Tailor your resume: Highlight Terraform, AWS, Kubernetes, and CI/CD experience. Emphasize AI usage in your work.
  • Showcase impact: Quantify achievements in infrastructure ownership, CI/CD reliability, and incident response.
  • Prepare for technical questions: Brush up on AWS services, Kubernetes, Terraform, and CI/CD best practices.
  • Discuss AI integration: Be ready to share examples of how you use AI tools in your workflow.
  • Ask insightful questions: Inquire about the platform's evolution, team culture, and AI's role.

Technical preparation

Master Terraform for IaC.,Deepen AWS knowledge: EKS, RDS, networking.,Practice Kubernetes production debugging.,Familiarize with CI/CD tools and AI.

Behavioral questions

Describe a complex incident you managed.,How do you balance developer velocity and reliability?,Share an example of using AI to improve work.,How do you influence infrastructure strategy?

Frequently asked questions

What is the salary range for a Staff Site Reliability Engineer at Babylist?
The starting salary range for this Staff Site Reliability Engineer role at Babylist is $226,673 to $271,991 annually. The final salary will depend on your location, experience, and qualifications, with potential for increases based on performance and role growth.
Is the Staff Site Reliability Engineer position at Babylist remote?
Yes, Babylist is a remote-first company, and this Staff Site Reliability Engineer position is open to team members across the U.S. and Canada. You can expect to work remotely, with occasional in-person gatherings twice a year.
What are the key technologies used by the Platform team at Babylist?
The Platform team at Babylist utilizes a tech stack including Ruby on Rails, AWS (with EKS, RDS, cloud networking), Sidekiq, MySQL, and Redis. They heavily rely on Terraform for Infrastructure as Code and various CI/CD tools like CircleCI or GitHub Actions.
How does Babylist incorporate AI into its engineering processes for SRE roles?
At Babylist, AI is not an experiment but an expectation. For the Staff SRE role, you are expected to naturally reach for AI in your work to move faster, improve output, explore tradeoffs, pressure-test designs, and generate code, allowing more focus on human judgment.
What are the primary responsibilities of a Staff SRE at Babylist?
As a Staff SRE, your primary responsibilities include owning and evolving AWS infrastructure using Terraform, ensuring CI/CD reliability, providing developer support, establishing monitoring and alerting standards, leading incident response, and contributing to platform strategy.
What kind of experience is Babylist looking for in a Staff Site Reliability Engineer candidate?
Babylist is looking for candidates with deep hands-on Terraform expertise, proven AWS experience at scale (EKS, RDS, networking), experience operating Kubernetes in production, comfort with CI/CD systems, strong observability instincts, and experience with on-call and incident management.
How does Babylist approach project failure for engineers?
Babylist has a culture where project failure and personal failure are not the same. They encourage rapid iteration and learning, and if something isn't working, they are willing to discard it and start over, fostering an environment where experimentation is valued.

Similar roles

Open positions we recommend based on this role.