PitchMeAI
Simple Life App

Senior Site Reliability Engineer

Simple Life App · EMEA

  • Hybrid
  • Full-time
  • $150,000 / year
  • EMEA
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Operate and manage AWS and Kubernetes infrastructure.
  • Automate processes using Go and infrastructure-as-code.
  • Ensure system reliability via CI/CD and observability.
  • Handle production incidents with strong debugging skills.
  • Collaborate across engineering, product, and security.

About the role

About Simple Life App

Simple Life is the #1 AI-powered health coaching app for adults who want to lose weight and enjoy a healthier lifestyle—without the stress or extremes. Our mission is to empower people to feel their best every day. By challenging traditional, restrictive approaches, Simple offers a more sustainable method grounded in ease, personalization, and real-life support. Simple has had over 17 million downloads and more than 300,000 5-star reviews, having helped millions lose weight successfully and sustainably. Simple has earned recognition as Best Virtual Coach and one of the Top 100 AI companies — all thanks to a dedicated global team driving real impact. With SIMPLE as a partner in their pocket, users feel cared for and empowered to embrace — and stick to — new healthy habits. To learn more, visit simple.life.

About the Role

Simple is looking for a Senior SRE to join our Platform team responsible for the AWS infrastructure, the Kubernetes platform and the internal tooling that the rest of engineering relies on. Push the pace of innovation and build a future of a healthier world with us!

This is an operations-led position. You will be working day to day on AWS, our infrastructure-as-code, our CI/CD setup, observability, and the on-call rotation. A meaningful part of the work is automation: when we find ourselves doing the same thing twice, we usually invest in tooling rather than writing another runbook. Most of that tooling is written in Go.

To give a sense of the environment: infrastructure is defined with Terraform and Terramate, with Atlantis running plan and apply on pull requests. Workloads run on EKS with Karpenter and Fargate, deployed through ArgoCD. Observability is built on Grafana, Loki, Tempo, and Prometheus compatible metrics.

We’re looking for

The most important quality for this role is how you handle problems that are not yet understood. Production incidents rarely present cleanly: logs can be incomplete, metrics can mislead, and the first plausible theory is often wrong. The right candidate stays focused under that kind of pressure, works through ambiguity in a structured way, and arrives at a real root cause rather than a convenient one. Strong investigation, debugging, and problem-solving instincts are essential. We also expect candidates to learn quickly. The stack and the company both move, and you will regularly be the first person on the team to take on something new.

  • Several years of hands-on experience operating production systems on AWS and Kubernetes, including genuine on-call ownership.
  • A solid working knowledge of AWS fundamentals, including VPC, IAM, EKS, and RDS.
  • Practical experience with Terraform and a GitOps-style delivery workflow (ArgoCD, Atlantis, Flux, or similar).
  • Comfort writing code, with some prior experience in Go or a willingness to pick it up (writing small services and tools is a regular part of the work).
  • Strong written and spoken English, and the communication skills to drive design discussions across engineering, product and security.

Perks and Benefits

  • Open-minded teams, a welcoming and inclusive company culture, plus the opportunity to make a real difference with a game-changing health tech product.
  • A competitive salary package based on your unique expertise, skillset, and impact on the product plus stock options.
  • In-office, remote and hybrid work opportunities.
  • The equipment whatever you need to be happy and productive.
  • A premium SIMPLE subscription.
  • 21 days annual leave, plus bank holidays (those observed where you live).
  • Flexible hours. We focus on your results, not how long you spend at your desk.

Key skills/competency

  • Site Reliability Engineering (SRE)
  • AWS
  • Kubernetes
  • Terraform
  • Go
  • CI/CD
  • Observability
  • Problem-solving
  • Debugging
  • Production Systems

Skills & topics

  • Site Reliability Engineer
  • SRE
  • AWS
  • Kubernetes
  • EKS
  • Terraform
  • Go
  • CI/CD
  • Observability
  • Infrastructure Automation
  • Cloud Engineering
  • DevOps

How to get hired

  • Tailor your resume: Highlight your AWS, Kubernetes, Terraform, and Go experience. Emphasize your problem-solving skills and on-call ownership.
  • Showcase automation: Detail your experience building tools and automating processes for CI/CD and observability.
  • Prepare for technical questions: Be ready to discuss AWS fundamentals, Kubernetes concepts, and debugging strategies for complex incidents.
  • Demonstrate problem-solving: Practice articulating how you approach ambiguous, high-pressure production issues and find root causes.
  • Research Simple Life: Understand their mission to empower healthier lifestyles and their AI-driven approach.

Technical preparation

Master AWS services like VPC, IAM, EKS, RDS.,Deepen Kubernetes knowledge (EKS, Karpenter, Fargate).,Practice Terraform and GitOps workflows.,Build projects using Go for automation tasks.

Behavioral questions

Describe handling an ambiguous production incident.,How do you stay focused under pressure?,Share an example of a new technology learned.,How do you drive design discussions across teams?

Frequently asked questions

What are the main responsibilities of a Senior Site Reliability Engineer at Simple Life App?
As a Senior SRE at Simple Life App, you'll manage AWS infrastructure, the Kubernetes platform, and internal tooling. Your day-to-day will involve working with AWS, infrastructure-as-code (Terraform), CI/CD, observability tools, and participating in an on-call rotation. A key focus is automating repetitive tasks with custom tooling, often in Go.
What technologies are used by the Platform team at Simple Life App for infrastructure and deployment?
The Platform team at Simple Life App utilizes Terraform and Terramate for infrastructure definition, with Atlantis for GitOps-style PR-based apply. Workloads run on EKS (Elastic Kubernetes Service) with Karpenter and Fargate, deployed via ArgoCD. Observability is managed using Grafana, Loki, Tempo, and Prometheus.
What kind of problems should I expect to solve as a Senior SRE at Simple Life App?
You should be prepared to tackle complex, ambiguous production incidents where logs and metrics might be incomplete or misleading. The role emphasizes structured problem-solving, debugging, and identifying true root causes under pressure. Learning new technologies and company processes quickly is also a key expectation.
Is prior experience with Go required for the Senior SRE role at Simple Life App?
While prior experience in Go is preferred, Simple Life App is also open to candidates willing to learn it. The role involves writing small services and tools in Go for automation, so a willingness to pick up the language is essential.
Does Simple Life App offer remote or hybrid work options for the Senior SRE position?
Yes, Simple Life App offers in-office, remote, and hybrid work opportunities for the Senior Site Reliability Engineer position, providing flexibility to suit your needs.