
DevOps/Observability Engineer
Quantiphi · United States
- Hybrid
- Full-time
- $150,000 / year
- United States
Job highlights
- Lead unified observability platform design and implementation.
- Architect observability pipeline on AWS.
- Expertise in OpenTelemetry and Kubernetes.
- Deploy Prometheus, Grafana, and Splunk.
- Hands-on technical leadership role.
About the role
About Quantiphi
Quantiphi is an award-winning, AI-First global digital engineering company that helps the world’s leading Fortune 1000 organizations transform bold ideas into measurable business impact. We go beyond building innovative AI technologies—we solve the problems that matter most to our clients.
Since our founding in 2013, Quantiphi has built a proven track record of turning complex challenges into meaningful outcomes across industries. Headquartered in Boston, with more than 4,000 professionals worldwide, we partner with global enterprises to deliver large-scale digital, cloud, and AI-driven transformation. #SolvingWhatMatters.
We are an Elite and Premier Partner to Google Cloud, AWS, NVIDIA, Snowflake, and other leading technology platforms, and our work has been recognized across the industry. We deliver First-in-class AI solutions across Life Sciences, Healthcare, Banking, Financial Services, CPG, Manufacturing, Energy, High-Tech, Telecommunications, etc., powered by cutting-edge Generative AI and Agentic AI accelerators. We are also proud to be certified as a Great Place to Work—reflecting our commitment to our people and our culture.
Role: Senior DevOps/Observability Engineer
We are seeking a highly experienced Senior DevOps/Observability Engineer with over 8 years of experience to lead the design and implementation of our next-generation, unified observability platform. This pivotal role will focus on architecting a sophisticated observability pipeline from the ground up, leveraging a modern, open-source-centric stack on Amazon Web Services (AWS). The ideal candidate will have deep expertise in designing and deploying observability solutions, with a strong emphasis on OpenTelemetry (OTel) and Kubernetes observability. You will be responsible for deploying, configuring, and integrating a suite of tools including Prometheus, Grafana, and Splunk to provide comprehensive insights into our complex, distributed systems. This is a hands-on role for a technical leader who is passionate about building scalable, reliable, and efficient monitoring and logging systems.
What You Will Do
- Design and implement end-to-end observability pipelines using OpenTelemetry, Prometheus, and Grafana on centralized infrastructure.
- Centralize AWS telemetry, including multi-account CloudTrail organization trails, cross-account CloudWatch metrics/logs, and VPC Flow Logs.
- Design log aggregation strategies, implement noise reduction/filtering at the collector level, and configure Splunk HTTP Event Collector (HEC) integrations.
- Build comprehensive alerting frameworks using Alertmanager and CloudWatch Alarms, coupled with advanced dashboard engineering in Grafana (using PromQL).
- Write Terraform modules for deploying and managing observability stacks and EC2 infrastructure.
- Manage, route, and optimize log pipelines at massive scale (TB/day).
- Deploy Prometheus and OTel within Kubernetes (EKS) or containerized (ECS) environments.
- Reduce observability spend through strategic metric dropping, log filtering, and efficient storage tiering.
What Is In It For You
- Join one of the world’s fastest-growing AI-first digital engineering companies and make a real impact at scale.
- Lead and collaborate with a high-energy team of talented, driven individuals solving complex, meaningful challenges.
- Work with Fortune 500 companies and disruptive innovators in a research-driven environment with 60+ patents.
- Stay ahead of the curve by gaining hands-on experience with cutting-edge AI, ML, data, and cloud technologies while continuously upskilling.
Key skills/competency
- DevOps
- Observability
- OpenTelemetry
- Kubernetes
- Prometheus
- Grafana
- Splunk
- AWS
- Terraform
- Log Management
Skills & topics
- DevOps Engineer
- Observability
- OpenTelemetry
- Prometheus
- Grafana
- Splunk
- AWS
- Kubernetes
- Terraform
- Cloud Engineering
- Monitoring
- Logging
- Site Reliability Engineer
- SRE
How to get hired
- Tailor your resume: Highlight experience with OpenTelemetry, Prometheus, Grafana, Splunk, and AWS IaC.
- Showcase expertise: Emphasize designing and implementing end-to-end observability pipelines.
- Quantiphi culture: Demonstrate your alignment with Quantiphi's AI-first, growth-oriented environment.
- Technical proficiency: Be ready to discuss specific AWS, Kubernetes, and log management strategies.
- Application process: Apply online and follow any specific instructions provided by Quantiphi.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the primary focus of the Senior DevOps/Observability Engineer role at Quantiphi?
- The primary focus is to lead the design and implementation of a next-generation, unified observability platform, architecting a sophisticated pipeline from scratch using a modern, open-source-centric stack on AWS.
- What are the key technologies involved in this DevOps role at Quantiphi?
- Key technologies include OpenTelemetry (OTel), Prometheus, Grafana, Splunk, Kubernetes (EKS), AWS services (CloudTrail, CloudWatch, VPC Flow Logs), and Infrastructure as Code (Terraform).
- Is this a remote position for the DevOps/Observability Engineer?
- Yes, this Senior DevOps/Observability Engineer position is fully remote within the USA.
- What level of experience is required for the DevOps/Observability Engineer at Quantiphi?
- A minimum of 8 years of experience is required for this Senior DevOps/Observability Engineer role.
- How does Quantiphi support continuous learning for its employees in this role?
- Quantiphi encourages continuous upskilling by providing hands-on experience with cutting-edge AI, ML, data, and cloud technologies.
- What is Quantiphi's partnership status with cloud providers like AWS?
- Quantiphi is an Elite and Premier Partner to AWS, along with other leading technology platforms like Google Cloud and NVIDIA.
- What is the expected impact of the Senior DevOps/Observability Engineer on Quantiphi's systems?
- The engineer will be responsible for providing comprehensive insights into complex, distributed systems by building scalable, reliable, and efficient monitoring and logging systems.
- How does Quantiphi address cost optimization in its observability solutions?
- The role involves optimizing observability spend through strategic metric dropping, log filtering, and efficient storage tiering, demonstrating a proven track record in cost reduction.
Similar roles
Open positions we recommend based on this role.