PitchMeAI
Juul Labs

Senior Site Reliability Engineer

Juul Labs · United States

  • Hybrid
  • Full-time
  • $185,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Own hybrid cloud infrastructure stability and performance.
  • Lead automation and architect for reliability.
  • Manage Nutanix, AWS, and GCP environments.
  • Handle critical incidents and L3 troubleshooting.
  • Develop infrastructure as code and optimize costs.

About the role

About Juul Labs

Juul Labs's mission is to transition the world’s billion adult smokers away from combustible cigarettes, eliminate their use, and combat underage usage of our products. We have the opportunity to address one of the world’s most intractable challenges through a commitment to exceptional quality, research, design, and innovation. Backed by leading technology investors, we are committed to the same excellence when it comes to hiring great talent. We are a diverse team that is united by this common purpose and we are hiring the world’s best engineers, scientists, designers, product managers, operations experts, and customer service and business professionals. If the opportunity to build your career is compelling, read on for more details.

Role and Responsibilities

A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance of Juul’s hybrid cloud infrastructure (Nutanix, AWS/GCP). This involves leading automation efforts, architecting for reliability, and acting as the final escalation point for critical incidents to ensure the platform is scalable and efficient.

Nutanix Platform Management

  • Design, deploy, and maintain enterprise-scale Nutanix AHV clusters and Prism Central for multi-cluster management.
  • Expert-level proficiency with Nutanix CLI (nCLI and acli) for advanced operations, troubleshooting, and automation.
  • Develop automation scripts using Nutanix REST APIs, Python SDK, PowerShell, and Terraform for infrastructure-as-code.
  • Create and manage VM templates, golden images, and standardized deployment catalogs for consistent provisioning.
  • Design disaster recovery solutions using Leap, Protection Domains, cross-cluster replication, and metro clustering.
  • Implement network micro-segmentation using Nutanix Flow and configure RBAC, encryption, and security hardening.
  • Lead L3 troubleshooting using advanced diagnostics, log analysis (CVM, Genesis), NCC health checks, and cluster service resolution.
  • Configure high availability, VM affinity rules, QoS policies, and optimize performance for mission-critical workloads.
  • Manage AHV networking with OVS bridges, VLANs, bonds, LACP and implement resource reservations and workload balance.

Hybrid Cloud Infrastructure Management

  • Design, deploy, and maintain hybrid cloud infrastructure across Nutanix HCI, AWS, and GCP platforms.
  • Architect and implement multi-cloud solutions ensuring high availability, scalability, and disaster recovery.

Cloud Platform Engineering

  • Architect and deploy enterprise-scale, highly available multi-cloud solutions across AWS and GCP with multi-region/multi-account strategies.
  • Expert-level proficiency with AWS CLI, GCP CLI, SDK, boto3, and Python for advanced automation and infrastructure orchestration.
  • Design AWS Organizations and GCP Organization hierarchies with consolidated billing, IAM policies, and centralized governance.
  • Configure and manage AWS Systems Manager (SSM) including Session Manager, Run Command, State Manager, and Automation for centralized fleet operations.
  • Implement centralized logging using CloudWatch/CloudTrail and GCP Cloud Logging with S3/Cloud Storage aggregation.
  • Integrate AWS and GCP with Splunk using HEC, CloudWatch subscriptions, Pub/Sub, Dataflow, and cloud-specific add-ons for SIEM correlation.
  • Design and deploy advanced load balancing solutions with AWS ALB/NLB/ELB and GCP Cloud Load Balancing including SSL termination and auto-scaling.
  • Develop infrastructure-as-code using Terraform, CloudFormation, CDK for repeatable multi-cloud deployments and CI/CD pipelines.
  • Configure AWS SSO, cross-account IAM roles, GCP Workload Identity, and federated access for centralized identity management.
  • Design VPC architectures with AWS Transit Gateway/PrivateLink and GCP Shared VPC/VPC peering for hybrid connectivity.
  • Manage containerized workloads using EKS, GKE, ECS, Cloud Run with service mesh, observability, and security best practices.
  • Implement disaster recovery using AWS Backup, Cross-Region Replication, GCP snapshots, and multi-region failover strategies.
  • Lead L3 troubleshooting using CloudWatch Insights, GCP Cloud Trace, VPC Flow Logs, X-Ray, and vendor support escalation.
  • Perform cost optimization through Reserved Instances, Committed Use Discounts, rightsizing, and automated resource lifecycle management.

System Administration

  • Administer and support Windows Server and Unix/Linux environments in production and non-production settings.
  • Perform OS-level hardening, patch management, and security compliance across heterogeneous systems.
  • Automate routine administrative tasks using PowerShell, Bash, Python, or similar scripting languages.
  • Manage GitHub organization settings, user permissions, repository access controls, and monitor GitHub Actions workflows and repository health across multiple teams.
  • Configure Splunk forwarders, heavy forwarders, and other integrations for data ingestion from cloud and on-premises sources.

Personal and Professional Qualifications

  • 8-12+ years infrastructure experience with 8+ years in Nutanix HCI and enterprise cloud (AWS/GCP).
  • Expert-level skills in Python, PowerShell, Bash scripting, infrastructure-as-code (Terraform/CloudFormation), and container orchestration (Kubernetes, EKS/GKE).
  • Proven experience managing enterprise-scale environments, hybrid cloud migrations, disaster recovery, and L3 critical incident management.
  • Strong networking knowledge (TCP/IP, VLANs, routing, VPN), security hardening, and compliance frameworks (ITIL).
  • Strategic thinker with exceptional analytical and troubleshooting abilities for complex multi-layer infrastructure issues.
  • Excellent communication skills to translate technical concepts to executives and non-technical stakeholders.
  • Calm under pressure during critical outages with meticulous attention to security, compliance, and configuration management.
  • Self-motivated continuous learner committed to staying current with evolving cloud technologies and automation opportunities.
  • Available for on-call rotations with strong documentation skills and customer service orientation.
  • Certifications (plus): Nutanix NCP/NCAP, AWS Solutions Architect Professional, AWS DevOps Professional, GCP Professional Cloud Architect, Terraform.

Education

Bachelor’s or master’s degree in computer science/IT.

Juul Labs Perks & Benefits

  • People. Work with talented, committed and supportive teammates.
  • Equity and performance bonuses. Every employee is a stakeholder in our success.
  • Cell phone subsidy, commuter benefits and discounts on JUUL products.
  • Excellent medical, dental and vision, disability, and life insurance, plus family support, wellness, legal, and employee assistance program benefits.
  • 401(k) plan with company matching.
  • Plus biannual discretionary performance bonuses.
Juul Labs is proud to be an equal opportunity employer and is committed to creating a diverse and inclusive work environment for all employees and job applicants, without regard to race, color, religion, sex, sexual orientation, age, gender identity or gender expression, national origin, disability or veteran status. We will consider for employment qualified applicants with arrest and conviction records, pursuant to the San Francisco Fair Chance Ordinance. Juul Labs also complies with the employment eligibility verification requirements of the Immigration and Nationality Act. All applicants must have authorization to work for Juul Labs in the US. LI-remote

Salary Ranges

Salary varies by role, level and location, and is dependent on the cost of labor in a given geographic region among other factors. These ranges may be modified at any time.

Locations

  • Tier 1 Locations: Greater New York City, and San Francisco Bay Area
  • Tier 2 Locations: Greater Boston, Washington DC Metropolitan Area, Seattle/Tacoma, Greater Sacramento, Southern California (Los Angeles/OC/San Diego, Riverside and Imperial counties)
  • Tier 3 Locations: Rest of New England, NY Capital District, Rest of New Jersey, Greater Philadelphia, Pittsburgh, Delaware, Rest of Maryland, Rest of Virginia, North Carolina, Atlanta, Miami-Fort Lauderdale-WPB, Chicagoland, Dallas, Houston, Austin, Minneapolis/St. Paul, Colorado, Phoenix, Las Vegas, Reno, Carson City NV., Portland Ore./Vancouver Wash., Rest of California, Hawaii
  • Tier 4 Locations: Rest of US including Alaska and Puerto Rico
  • Tier 1 Range: $185,000 USD - $227,000 USD
  • Tier 2 Range: $168,000 USD - $206,000 USD
  • Tier 3 Range: $158,000 USD - $194,000 USD
  • Tier 4 Range: $141,000 USD - $173,000 USD

Key skills/competency

  • Site Reliability Engineering
  • Nutanix
  • AWS
  • GCP
  • Python
  • Terraform
  • Kubernetes
  • CI/CD
  • Disaster Recovery
  • Incident Management

Skills & topics

  • Senior Site Reliability Engineer
  • SRE
  • Nutanix
  • AWS
  • GCP
  • Python
  • Terraform
  • Kubernetes
  • Cloud Infrastructure
  • Automation
  • Hybrid Cloud
  • DevOps
  • Infrastructure as Code
  • Disaster Recovery
  • Incident Management
  • System Administration
  • Linux
  • Windows Server
  • Splunk
  • GitHub

How to get hired

  • Tailor your resume: Highlight Nutanix, AWS, GCP, Python, Terraform, and Kubernetes experience.
  • Showcase automation skills: Provide examples of scripting and infrastructure-as-code projects.
  • Demonstrate problem-solving: Detail experience with L3 troubleshooting and incident management.
  • Prepare for technical questions: Be ready to discuss hybrid cloud architecture and DR strategies.
  • Research Juul Labs: Understand their mission and values to align your application.

Technical preparation

Master Nutanix CLI, APIs, and Python SDK.,Deep dive into AWS and GCP services.,Practice Terraform and CloudFormation scripting.,Prepare Kubernetes and container orchestration scenarios.

Behavioral questions

Describe a critical incident you resolved.,How do you approach infrastructure automation?,Explain a complex multi-cloud design challenge.,How do you handle pressure during outages?

Frequently asked questions

What are the key responsibilities for a Senior Site Reliability Engineer at Juul Labs?
The Senior Site Reliability Engineer at Juul Labs is responsible for owning the operational stability and performance of their hybrid cloud infrastructure (Nutanix, AWS/GCP). This includes leading automation initiatives, architecting for reliability, and acting as the final escalation point for critical incidents to ensure a scalable and efficient platform.
What is Juul Labs' mission and how does it relate to this role?
Juul Labs' mission is to transition adult smokers away from combustible cigarettes. This mission drives a need for highly reliable, scalable, and secure technology infrastructure, which is the core responsibility of the Senior Site Reliability Engineer.
What are the primary cloud platforms managed by the SRE at Juul Labs?
The Senior Site Reliability Engineer will manage Juul Labs' hybrid cloud infrastructure, which includes Nutanix HCI, Amazon Web Services (AWS), and Google Cloud Platform (GCP).
What technical skills are most important for this Senior SRE position?
Expert-level skills in Python, PowerShell, Bash scripting, infrastructure-as-code (Terraform/CloudFormation), and container orchestration (Kubernetes, EKS/GKE) are crucial. Strong knowledge of Nutanix, AWS, and GCP is also essential.
What kind of experience is required for the Senior Site Reliability Engineer role?
The role requires 8-12+ years of infrastructure experience, with at least 8 years specifically in Nutanix HCI and enterprise cloud environments (AWS/GCP). Proven experience in managing enterprise-scale environments, hybrid cloud migrations, disaster recovery, and L3 critical incident management is also necessary.
What are the benefits of working at Juul Labs as a Senior SRE?
Juul Labs offers competitive benefits including equity and performance bonuses, a cell phone subsidy, commuter benefits, discounts on products, comprehensive medical, dental, and vision insurance, disability and life insurance, family support, wellness programs, legal services, employee assistance programs, and a 401(k) plan with company matching.
Are there specific certifications that are beneficial for this role at Juul Labs?
While not strictly required, certifications such as Nutanix NCP/NCAP, AWS Solutions Architect Professional, AWS DevOps Professional, GCP Professional Cloud Architect, and Terraform certifications are considered a plus.
What is the expected work arrangement for this Senior Site Reliability Engineer role?
The job posting indicates 'LI-remote', suggesting that this position can be performed remotely.