
Staff Production Operations Engineer
Redpanda Data · United States
- Hybrid
- Full-time
- $256,000 / year
- United States
Job highlights
- Drive reliability operations and process improvements.
- Coordinate on-call and incident management globally.
- Build AI agents to automate operational tasks.
- Maintain runbooks and incident process documentation.
- Requires 5+ years in SRE/DevOps large-scale environments.
About the role
About Redpanda
Redpanda is pioneering the Agentic Data Plane (ADP) - a new category in AI infrastructure that makes it simple and secure to connect AI agents with enterprise data and systems. Built on a multi-modal data streaming engine, Redpanda empowers agentic applications that reason and act in real-time with speed, autonomy, and precision.
Global leaders including Activision Blizzard, Cisco, Moody's, Texas Instruments, Vodafone and 2 of the top 5 banks in the U.S. rely on Redpanda to process hundreds of terabytes of data a day.
Backed by premier venture investors Lightspeed, GV and Haystack VC, Redpanda is a diverse, people-first organization with teams distributed around the globe.
About the Role
We're looking for a Staff Production Operations Engineer to drive Redpanda's reliability operations program. This role combines hands-on site reliability engineering with planning and coordination skills to ensure a world-class operations practice across a globally distributed engineering team.
In this role, you'll work with the broader Engineering team, Engineering leadership, Product and Customer Success to drive operational excellence. You'll coordinate our on-call and incident lead rotations, drive blameless post-incident reviews, and own the processes that help us respond faster, learn more from outages, and systematically improve reliability. We're looking for someone who can leverage AI agents to automate the operational toil that slows teams down, building on Redpanda's own ADP platform to do it.
What You Will Do:
- Drive process improvements across the incident lifecycle: severity models, triage enforcement, alert noise reduction, and follow-up completion rates
- Coordinate the on-call program across multiple geographies: manage schedules and shadow rotations, onboard new engineers, and ensure consistent coverage
- Select incidents for post-incident review, facilitate blameless post-incident reviews, document findings, and track follow-up completion. Contribute to addressing incident follow-ups where possible, either by fixing issues directly or prototyping solutions
- Build AI agents to automate operational toil, including oncall automation, as well as incident summarization, post-incident reviews prep, follow-up tracking, and on-call analytics
- Maintain runbooks, playbooks, and incident process documentation, and keep them current as processes evolve
What You Bring:
- 5+ years of experience in site reliability engineering, DevOps, or production operations in large-scale, highly reliable environments
- A track record of leading initiatives end-to-end, from design and planning, to execution and production operation
- Hands-on experience with incident management tooling (incident.io, PagerDuty, or similar) and observability stacks (Datadog, Grafana, Sentry, CloudWatch, or equivalent)
- Strong Fluency with reliability concepts: MTTD, MTTR, MTTA, error budgets, SLOs
- Experience building automation and tooling to reduce operational toil
- Proficiency in Go (or comparable systems language with willingness to ramp)
- Experience with AI-assisted software development workflows including tools like Claude Code
- Working knowledge of at least one of AWS / Azure / GCP, including infrastructure as code for system and network infrastructure
- Strong written communication; ability to drive alignment across engineering teams without direct authority
Nice to Have:
- Hands-on experience building agents or automations using LLMs
- Familiarity with Redpanda, Apache Kafka, or other streaming infrastructure
- Prior experience in a fast-growing B2B infrastructure or developer tools company
Compensation
U.S. base salary range for this role is $220,000 - $256,000 (CA, NY, WA) and $211,000 - $250,000 (other US locations). Our salary ranges are determined by role, level, and location. We strive to consider each candidate's job-related skills, location, experience, relevant education or training to determine individual base salary. Your talent partner will share more about the specific salary range for your preferred location during the hiring process.
AI in Hiring
Please note that Redpanda uses artificial intelligence (AI) technology to assist in the screening and assessment of applications for this position. However, all final hiring decisions are made by our human hiring team.
Vacancy Status
This job posting is for an existing vacancy.
Join Us
Join Redpanda if you’d enjoy being part of a fast-moving, diverse, people-first organization with team members around the globe and a culture based on trust, transparency, communication, and kindness. You'll dive into a nimble, high-impact team with the latest AI tools — and the budget to actually use them.
Key skills/competency
- Staff Production Operations Engineer
- Site Reliability Engineering (SRE)
- DevOps
- Production Operations
- Incident Management
- Observability
- Automation
- Go Programming
- Cloud Infrastructure (AWS/Azure/GCP)
- AI Agents
Skills & topics
- Staff Production Operations Engineer
- Site Reliability Engineering
- DevOps
- Production Operations
- Incident Management
- Incident Response
- Reliability Engineering
- Go
- AWS
- Azure
- GCP
- Automation
- AI Agents
- LLM
How to get hired
- Tailor your resume: Highlight 5+ years of SRE/DevOps experience, incident management, automation, and proficiency in Go. Emphasize leadership in end-to-end initiatives.
- Showcase AI/LLM skills: Detail any experience building AI agents or using AI-assisted development workflows. Mention familiarity with tools like Claude Code.
- Demonstrate reliability expertise: Clearly articulate your understanding of reliability concepts like MTTD, MTTR, and SLOs. Include experience with observability stacks and incident management tooling.
- Highlight cloud and IaC knowledge: Specify your working knowledge of AWS, Azure, or GCP, and experience with infrastructure as code.
- Prepare for technical and behavioral interviews: Be ready to discuss your experience with Go, large-scale systems, and how you drive alignment across teams.
Technical preparation
Behavioral questions
Frequently asked questions
- What is Redpanda's approach to AI in the hiring process for a Staff Production Operations Engineer?
- Redpanda utilizes AI technology to aid in the screening and assessment of applications for the Staff Production Operations Engineer role. However, all final hiring decisions are exclusively made by their human hiring team, ensuring a balanced approach to candidate evaluation.
- How does Redpanda ensure consistent on-call coverage across multiple geographies for this role?
- The Staff Production Operations Engineer will coordinate the on-call program across multiple geographies. This includes managing schedules, shadow rotations, onboarding new engineers, and ensuring consistent coverage to maintain operational excellence.
- What specific AI technologies does Redpanda expect a Staff Production Operations Engineer to leverage?
- The role involves building AI agents to automate operational toil, such as on-call automation, incident summarization, post-incident review preparation, follow-up tracking, and on-call analytics. Experience with AI-assisted software development workflows and tools like Claude Code is also a plus.
- What are the key reliability concepts a Staff Production Operations Engineer at Redpanda should be fluent in?
- A strong fluency in reliability concepts is crucial. This includes understanding and applying metrics like Mean Time To Detect (MTTD), Mean Time To Resolve (MTTR), Mean Time To Acknowledge (MTTA), error budgets, and Service Level Objectives (SLOs).
- How does Redpanda handle compensation for the Staff Production Operations Engineer role?
- Redpanda provides U.S. base salary ranges that vary by location. For example, ranges are $220,000 - $256,000 for CA, NY, WA, and $211,000 - $250,000 for other US locations. The final salary is determined by role, level, location, skills, experience, and education.
- What is the expected experience level for a Staff Production Operations Engineer at Redpanda?
- Candidates should have at least 5 years of experience in site reliability engineering, DevOps, or production operations within large-scale, highly reliable environments. A proven track record of leading initiatives end-to-end is also essential.
- Does Redpanda have a preference for specific cloud providers for the Staff Production Operations Engineer position?
- The role requires working knowledge of at least one of AWS, Azure, or GCP. Experience with infrastructure as code for system and network infrastructure within these cloud environments is also expected.