
Senior DevOps Engineer
TrueML · United States
- Hybrid
- Full-time
- $150,000 / year
- United States
Job highlights
- Lead cloud architecture and CI/CD pipelines.
- Design self-service internal developer platforms.
- Optimize AWS cloud spend and infrastructure.
- Build high availability and disaster recovery systems.
- Implement advanced monitoring and AIOps.
About the role
About the Role
TrueML Products is seeking a highly experienced Sr. DevOps Engineer I to serve as a core contributor on our infrastructure and platform engineering efforts. This role is critical in execution-focused cloud architecture, establishing robust CI/CD pipelines, and ensuring the absolute scalability, security, and reliability of our products. Reporting to the Sr. Manager, DevOps, you will drive the day-to-day evolution of our internal developer platform and infrastructure-as-code (IaC) architecture. The ideal candidate is a deeply technical, hands-on engineer with a "systems-thinking" mindset. We are looking for a practitioner who thrives on solving complex distributed systems challenges and considers leveraging GenAI and AIOps tooling second-nature for optimizing system performance, monitoring, and automation.
What You'll Do (Technical Execution & Architecture)
- Implement the technical roadmap for Infrastructure as Code (IaC), CI/CD evolution, and cloud-native architecture to support TrueML’s scaling needs.
- Design, develop, and maintain self-service internal platforms to reduce developer cognitive load, enabling feature teams to deploy and manage services with minimal friction at increased velocity.
- Act as a core steward for cloud spend (AWS), proactively identifying and driving cost-optimization initiatives across our infrastructure.
- Build and maintain infrastructure architecture that supports strict High Availability (HA) requirements and robust Disaster Recovery (DR) protocols across multiple regions.
- Implement and evolve comprehensive monitoring, logging, and distributed tracing systems, leveraging AIOps to move from reactive to predictive system maintenance.
What You'll Do (Deep-dive Hands-On Engineering)
- Write and review high-quality, production-grade code in languages like Python, Go, or Bash to automate complex operational tasks and system integrations.
- Drive hands-on development of robust Terraform Infrastructure as Code for reliable resource provisioning.
- Directly architect, optimize, and troubleshoot complex CI/CD workflows (GitHub Actions, ArgoCD, Atlantis) to maximize build-and-deploy speed and reliability.
- Proactively manage, fine-tune, and scale container orchestration environments, including hands-on configuration of Ingress controllers and declarative GitOps workflows.
- Manage the technical integration and API configurations between various tools in the DevOps stack (e.g., connecting Jira, VictorOps, Slack, and Observe for seamless incident flow).
What You'll Do (Collaboration & Knowledge Sharing)
- Partner closely with other Senior DevOps Engineers and Engineering Managers to align infrastructure deliverables with product roadmaps, ensuring DevOps acts as an accelerator.
- Collaborate with Quality Engineering and Security teams to enforce "Definition of Done" standards that include automated testing and security gates.
- Provide technical guidance to junior engineers on the team, fostering a culture of continuous learning.
Who You Are:
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 6+ years of experience in DevOps, Site Reliability Engineering (SRE), or Software Engineering, working within high-performing senior engineering teams.
- Expert-level mastery with AWS and hands-on experience managing multi-region, high-availability deployments.
- Advanced experience with Kubernetes (K8s) and Docker, including cluster management, networking, and scaling in production environments.
- High proficiency in Terraform to drive consistency and automation across all infrastructure layers (Experience with Atlantis is a plus).
- Deep experience designing and maintaining complex pipelines (GitHub Actions, GitLab CI, or Jenkins) and mastery of scripting languages like Python, Go, or Bash.
- Hands-on experience with modern monitoring, observability, and tracing stacks (Datadog, Observe) and a firm grasp of SRE principles (SLIs/SLOs/Error Budgets).
- Experience acting as an Incident Commander or critical responder for high-severity outages.
- Experience integrating AI-assisted productivity tools (Cline, GitHub Copilot) into your engineering workflow to accelerate delivery, troubleshooting, and system monitoring.
Key skills/competency
- DevOps
- Site Reliability Engineering (SRE)
- Cloud Architecture
- CI/CD
- Infrastructure as Code (IaC)
- AWS
- Kubernetes (K8s)
- Terraform
- Python
- Monitoring and Observability
Skills & topics
- Senior DevOps Engineer
- DevOps
- Site Reliability Engineering
- SRE
- Cloud Architecture
- CI/CD
- Infrastructure as Code
- IaC
- AWS
- Kubernetes
- K8s
- Docker
- Terraform
- Python
- Go
- Bash
- Monitoring
- Observability
- AIOps
- GitHub Actions
- ArgoCD
- Atlantis
- Datadog
- High Availability
- Disaster Recovery
How to get hired
- Tailor your resume: Highlight your 6+ years of DevOps/SRE experience, AWS mastery, Kubernetes, Terraform, and CI/CD expertise.
- Showcase coding skills: Emphasize your proficiency in Python, Go, or Bash and experience with incident response.
- Demonstrate IaC proficiency: Detail your experience with Terraform and tools like GitHub Actions, ArgoCD, or Atlantis.
- Prepare for technical deep-dives: Be ready to discuss distributed systems, container orchestration, and AIOps strategies.
- Highlight AI tool experience: Mention your use of AI-assisted tools like GitHub Copilot in your workflow.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key technical skills for the Senior DevOps Engineer role at TrueML Products?
- The Senior DevOps Engineer role at TrueML Products requires expert-level mastery of AWS, advanced experience with Kubernetes and Docker, high proficiency in Terraform, and deep experience with CI/CD pipelines (GitHub Actions, ArgoCD, Atlantis). Strong scripting skills in Python, Go, or Bash, along with experience in modern monitoring/observability stacks (Datadog, Observe) and SRE principles, are also crucial.
- What kind of experience is expected for the Senior DevOps Engineer position at TrueML Products?
- TrueML Products is looking for candidates with 6+ years of experience in DevOps, SRE, or Software Engineering, specifically within high-performing senior engineering teams. Experience managing multi-region, high-availability deployments on AWS, acting as an Incident Commander, and integrating AI-assisted productivity tools are highly valued.
- How does TrueML Products leverage AI in its DevOps practices?
- TrueML Products actively integrates GenAI and AIOps tooling to optimize system performance, monitoring, and automation. The ideal candidate will find leveraging these tools second-nature and should have experience using AI-assisted productivity tools like Cline or GitHub Copilot to accelerate delivery and troubleshooting.
- What is the role of Infrastructure as Code (IaC) at TrueML Products?
- Infrastructure as Code (IaC) is a core component of TrueML Products' strategy. The Senior DevOps Engineer will implement the technical roadmap for IaC using tools like Terraform to ensure reliable resource provisioning, consistency, and automation across all infrastructure layers, supporting the company's scaling needs.
- What are the collaboration expectations for a Senior DevOps Engineer at TrueML Products?
- Senior DevOps Engineers at TrueML Products are expected to partner closely with other senior engineers and engineering managers to align infrastructure with product roadmaps. They also collaborate with Quality Engineering and Security teams to enforce 'Definition of Done' standards and provide technical guidance to junior engineers, fostering a culture of continuous learning.