
Staff Site Reliability Engineer, Production Engineering
Dropbox · United States
- Hybrid
- Full-time
- $268,600 / year
- United States
Job highlights
- Define company-wide reliability strategy for AI-driven development.
- Lead initiatives to reduce reliability risks at scale.
- Partner with teams on monitoring, alerting, and incident response.
- Identify and mitigate AI-related development risks.
- Provide technical leadership and mentorship to engineers.
About the role
About the Role
As a Site Reliability Engineer at Dropbox, you will be instrumental in shaping the company's reliability strategy, especially as AI technologies transform software development and operations. This role is crucial for enhancing stability, observability, incident response, and overall operational excellence. You will define how reliability adapts to agentic development and AI-enabled software delivery, preparing Dropbox for increased complexity, higher pull request volumes, and evolving incident patterns. Your work will involve close collaboration with Engineering, Product, and leadership teams to elevate reliability standards, influence long-term platform investments, and ensure millions of users continue to have a dependable experience.
Dropbox operates on a Virtual First model, and this position is open to candidates in Zones 2 and 3 within the US. For details on specific neighborhoods within these zones, please refer to the Compensation section.
Responsibilities
- Define and advance Dropbox’s company-wide technical reliability strategy to align with AI-assisted and agentic software development paradigms.
- Establish multi-year reliability objectives, standards, and roadmaps covering observability, debugging, incident management, service health, and operational readiness.
- Spearhead cross-team initiatives to mitigate reliability risks associated with increased software delivery velocity, pull request volume, service complexity, and incident frequency.
- Collaborate with engineering leaders and platform teams to enhance monitoring, alerting, debugging, SLOs, SLAs, and company-wide incident response systems.
- Identify and address emerging reliability risks stemming from AI-enabled development, designing scalable systems and guardrails for mitigation.
- Provide technical leadership and mentorship to engineers, promoting high engineering quality, sound reliability judgment, and operational excellence.
- Ensure clear communication and alignment with senior stakeholders on reliability priorities, tradeoffs, risks, and progress.
On-Call Rotation
Many teams at Dropbox participate in on-call rotations, requiring availability during both core and non-core business hours. Participation in these rotations is a standard part of employment. Candidates are encouraged to inquire about specific rotation details for the team they are applying to.
Requirements
- BS degree in Computer Science or a related technical field (e.g., physics, mathematics) or equivalent practical experience.
- 12+ years of experience in software engineering, site reliability engineering, infrastructure engineering, or related technical fields.
- Proven track record of defining and executing multi-year, multi-team strategies for reliability, infrastructure, or platforms with significant business and customer impact.
- Extensive experience with distributed systems, production operations, observability, incident response, SLOs/SLAs, debugging, and reliability risk management.
- Demonstrated ability to diagnose complex technical issues, debug production systems, automate operational workflows, and design resilient software components.
- Experience influencing engineering roadmaps across multiple teams and making strategic technical decisions for the broader organization.
- Excellent communication and collaboration skills, with the ability to navigate ambiguity and drive execution across diverse teams.
Preferred Qualifications
- Experience adapting reliability strategies, developer tooling, or operational processes for AI-assisted software development.
- Experience building or scaling observability, debugging, incident management, or developer productivity platforms for large organizations.
- Experience leading reliability enhancements in high-deployment-velocity environments with complex service dependencies.
- A history of mentoring senior engineers, establishing technical standards, and disseminating reliability best practices.
- Familiarity with AI-enabled tooling, agentic development, or operational risks associated with rapid automation in software development.
Compensation
US Zone 2: $223,400—$302,200 USD
US Zone 3: $198,600—$268,600 USD
Note: Salary is one component of a total rewards package, which may include corporate bonuses and stock (RSUs). Specific pay is determined by factors such as job level, location, skills, and peer compensation. Offers typically fall between the minimum and midpoint of the stated range.
Key skills/competency
- Staff Site Reliability Engineer
- Reliability Strategy
- AI-Assisted Development
- Observability
- Incident Response
- Distributed Systems
- Production Operations
- SLOs and SLAs
- Technical Leadership
- Cross-functional Collaboration
Skills & topics
- Site Reliability Engineer
- SRE
- Staff Engineer
- Production Engineering
- Reliability
- Observability
- Incident Response
- Distributed Systems
- Cloud Infrastructure
- AI in Software Development
- Python
- Go
- Kubernetes
- AWS
- Technical Leadership
How to get hired
- Tailor your resume: Highlight your 12+ years of experience in SRE, distributed systems, and defining multi-year reliability strategies. Emphasize your ability to influence roadmaps and lead cross-functional initiatives.
- Showcase AI expertise: Detail any experience you have adapting reliability practices or tooling for AI-assisted development workflows.
- Prepare for technical interviews: Be ready to discuss complex distributed systems, production operations, debugging production issues, and designing resilient software.
- Demonstrate leadership: Prepare examples of how you've provided technical leadership, mentored engineers, and driven alignment among senior stakeholders.
- Research Dropbox's culture: Understand their Virtual First model, AI fluency principles, and commitment to enlightened work practices.
Technical preparation
Behavioral questions
Frequently asked questions
- What is Dropbox's Virtual First policy for the Staff Site Reliability Engineer role?
- Dropbox operates under a Virtual First model, meaning this Staff Site Reliability Engineer role is primarily remote. However, this position requires approximately 5-10% travel for team gatherings and offsites. Candidates must reside in US Zones 2 or 3.
- How does Dropbox incorporate AI into the Staff Site Reliability Engineer role?
- AI fluency is a core behavior at Dropbox. For this Staff Site Reliability Engineer position, you will help define reliability strategies for AI-assisted development, identify AI-related risks, and leverage AI tools to improve workflows and enhance impact. Ownership, experimentation, leverage, and learning are key AI behaviors expected.
- What are the salary expectations for a Staff Site Reliability Engineer at Dropbox?
- The expected annual salary range for this role depends on the US Zone: $223,400–$302,200 USD for Zone 2, and $198,600–$268,600 USD for Zone 3. The final offer will consider factors like location, skills, and experience.
- What technical experience is required for the Staff Site Reliability Engineer position at Dropbox?
- The Staff Site Reliability Engineer role requires a BS degree in CS or equivalent experience, plus 12+ years in SRE or related fields. Key technical areas include deep experience with distributed systems, production operations, observability, incident response, SLOs/SLAs, debugging, and reliability risk management.
- Does this Staff Site Reliability Engineer role involve on-call duties at Dropbox?
- Yes, many teams at Dropbox have on-call rotations, and participation is expected as part of the Staff Site Reliability Engineer role. Candidates are encouraged to ask for specific details regarding the rotation for the team they are applying to during the interview process.
- What is considered 'AI Fluency' at Dropbox for this role?
- AI Fluency at Dropbox means thoughtfully and effectively using AI tools to improve your work and support others, rather than requiring deep expertise. For the Staff Site Reliability Engineer, this involves responsible ownership, exploring new AI capabilities, leveraging AI for efficiency, and continuously learning about AI trends.
- How does Dropbox handle compensation for remote employees in different zones?
- Dropbox determines compensation for remote employees based on the zip code of their work location, which then assigns them to a specific US Zone (Zone 1, 2, or 3). Each zone has a corresponding salary range for the Staff Site Reliability Engineer role.