
Sr. Site Reliability Engineer
Tavily · New York, United States
- On site
- Full-time
- $150,000 / year
- New York, United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Manage Kubernetes and infrastructure as code.
- Optimize real-time data pipelines and CI/CD.
- Build monitoring, alerting, and observability stacks.
- Debug production issues and manage cloud costs.
- Own the entire infrastructure for an AI company.
About the role
About Tavily
We're building the infrastructure layer for agentic web interaction at scale. Our API is designed from the ground up to power Retrieval-Augmented Generation (RAG) and real-time reasoning in AI systems. By connecting LLMs to high-quality, trustworthy web content, we help developers build agents that are not only intelligent — but also informed.
We work with some of the most innovative teams in AI — from small startups shaping the ecosystem to the largest enterprises deploying AI at scale. Whether it's powering sales assistants, research copilots, or internal knowledge tools, we're the missing link between LLMs and the real world.
The Role: Senior Site Reliability Engineer
- Managing Kubernetes clusters across multiple environments and regions
- Owning infrastructure as code for all resources
- Maintaining and improving CI/CD pipelines and GitOps-based deployments
- Maintaining and optimize real-time data pipelines that process billions of events per day across distributed queues and stream processors
- Building out monitoring, alerting, and observability
- Debugging production issues across services
- Managing cloud costs and capacity planning
- Working closely with a small engineering team — you'd own infra, not a slice of it
What we're looking for
- 5-8 years in a DevOps or SRE role, working in production environments
- Proven experience designing and operating large-scale, distributed systems, with a solid understanding of API design, reliability, and performance at scale
- Strong Kubernetes experience in a managed cloud environment
- Proficiency with infrastructure as code (Terraform or similar)
- Experience with GitOps-based deployment workflows
- Built or maintained observability stacks (logging, metrics, alerting)
- Experience handling production incidents calmly and methodically
Nice to have:
- Multi-region deployments
- Search infrastructure
- Data pipeline experience (streaming, warehousing)
- Proxy/networking infrastructure at scale
Why Tavily?
- Full ownership — small team, you own the entire infrastructure, not a slice of it
- Real scaling challenges — bursty scraping workloads, cache invalidation, multi-region, millions of daily requests
- AI-native company — your infra directly powers AI agents used by leading companies in the space.
Key employee benefits in the US:
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Key skills/competency
- Site Reliability Engineering
- Kubernetes
- Infrastructure as Code
- CI/CD
- GitOps
- Observability
- Distributed Systems
- Production Incidents
- Cloud Cost Management
- Capacity Planning
Skills & topics
- Site Reliability Engineer
- SRE
- DevOps
- Kubernetes
- Infrastructure as Code
- Terraform
- GitOps
- CI/CD
- Observability
- Cloud
- AWS
- GCP
- Azure
- Distributed Systems
- Production Support
- AI Infrastructure
- Data Pipelines
- Monitoring
- Alerting
- Capacity Planning
How to get hired
- Tailor your resume: Highlight your 5-8 years of SRE/DevOps experience and specific achievements in large-scale distributed systems, Kubernetes, and infrastructure as code (Terraform).
- Showcase GitOps & Observability: Emphasize your experience with GitOps workflows and building/maintaining observability stacks (logging, metrics, alerting) in production environments.
- Demonstrate incident management skills: Prepare to discuss how you've calmly and methodically handled production incidents and managed cloud costs.
- Research Tavily's AI focus: Understand their mission of powering AI agents and how your infra skills contribute to their success.
Technical preparation
Practice managing Kubernetes clusters.,Build projects with Terraform.,Set up CI/CD and GitOps pipelines.,Implement monitoring and alerting systems.
Behavioral questions
Describe a challenging production incident.,How do you prioritize infrastructure tasks?,How do you handle debugging complex systems?,How do you manage cloud costs effectively?
Frequently asked questions
- What is Tavily's mission and what kind of infrastructure challenges will a Senior Site Reliability Engineer face?
- Tavily's mission is to build the infrastructure layer for agentic web interaction at scale, connecting LLMs to trustworthy web content. As a Senior Site Reliability Engineer, you will face significant scaling challenges, including managing Kubernetes clusters across multiple environments, optimizing real-time data pipelines processing billions of events daily, and ensuring the reliability of systems powering AI agents for leading companies.
- What are the key technical requirements for the Senior Site Reliability Engineer role at Tavily?
- The key technical requirements include 5-8 years of experience in DevOps or SRE, strong Kubernetes expertise in a managed cloud environment, proficiency with infrastructure as code (like Terraform), experience with GitOps deployments, and a proven track record of building and maintaining observability stacks (logging, metrics, alerting).
- What does 'full ownership' mean for the Senior Site Reliability Engineer at Tavily?
- Full ownership at Tavily means you will be responsible for the entire infrastructure of the company, not just a specific slice. Given the small engineering team, you will have a significant impact and autonomy in managing and evolving the infrastructure that powers Tavily's AI-native solutions.
- What are the benefits of working at Tavily as a Senior Site Reliability Engineer, especially regarding work-life balance and remote work?
- Tavily offers comprehensive benefits including 100% company-paid health insurance for employees and families, a 401(k) plan with a company match, generous parental leave, and up to $85/month reimbursement for remote work expenses like mobile and internet. The role is remote.
- How does Tavily's focus on AI impact the Senior Site Reliability Engineer role?
- As an AI-native company, the infrastructure you manage directly powers AI agents used by leading companies. This means you'll be working on cutting-edge technology and contributing to the core of AI development, dealing with unique challenges related to AI workloads and data processing at scale.
Similar roles
Open positions we recommend based on this role.