
Site Reliability Engineer
hackajob · United States
- Hybrid
- Full-time
- $120,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Ensure AI system reliability and stability.
- Manage incident response and root cause analysis.
- Optimize observability with modern tools.
- Oversee capacity, performance, and FinOps.
- Monitor SLOs and SLAs for critical systems.
About the role
Site Reliability Engineer - MANTECH Initiative
hackajob is collaborating with MANTECH to connect them with exceptional professionals for this role. MANTECH seeks motivated, career, and customer-oriented Site Reliability Engineers (SREs) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-scale AI, data, and mission platform capabilities across cloud, edge, and classified operational environments.
This role supports the operational reliability, scalability, monitoring, and incident response for the enterprise AI systems. You will focus on operational outcomes and optimizing system performance.
Responsibilities Include But Are Not Limited To
- Apply core reliability engineering principles to ensure high availability and stability of production systems.
- Manage incident response, root cause analysis, and post-mortem processes for the AI platform.
- Implement and optimize observability operations using OpenTelemetry, Prometheus, Grafana, Loki, or Tempo.
- Oversee capacity planning, performance optimization, and FinOps practices.
- Define and continuously monitor Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
Minimum Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 5 or more years of experience in Site Reliability Engineering (SRE), DevOps, or production operations.
- Extensive experience with cloud-native infrastructure, particularly Kubernetes.
- Deep knowledge of monitoring, alerting, and logging systems.
- Proven ability to automate operational tasks and reduce toil.
Preferred Qualifications
- Hands-on experience with the full observability stack: OpenTelemetry, Prometheus, Grafana, Loki, and Tempo.
- Experience with FinOps and optimizing cloud resource consumption.
- Experience supporting high-scale distributed systems in a secure environment.
Clearance Requirements
For onsite work, a TS/SCI clearance with Poly will be required.
Physical Requirements
- The person in this position must be able to remain in a stationary position 50% of the time.
- Frequently communicates with co-workers, management, and customers, which may involve delivering presentations.
- Constantly operates a computer and other office productivity machinery.
Key skills/competency
- Site Reliability Engineering (SRE)
- DevOps
- Kubernetes
- Observability
- OpenTelemetry
- Prometheus
- Grafana
- FinOps
- Capacity Planning
- Incident Response
Skills & topics
- Site Reliability Engineer
- SRE
- DevOps
- Kubernetes
- Cloud
- AI
- Observability
- Prometheus
- Grafana
- FinOps
How to get hired
- Tailor your resume: Highlight SRE experience, Kubernetes, and observability tools relevant to MANTECH.
- Showcase automation skills: Provide examples of reducing toil and improving operational efficiency.
- Demonstrate cloud-native expertise: Emphasize experience with cloud platforms and container orchestration.
- Prepare for technical questions: Be ready to discuss incident response, monitoring, and system design.
- Understand MANTECH's mission: Research their work with AI and defense initiatives.
Technical preparation
Master Kubernetes and cloud-native architecture.,Deep dive into observability tools: Prometheus, Grafana, OpenTelemetry.,Practice incident response and root cause analysis scenarios.,Automate common operational tasks with scripting.
Behavioral questions
Describe a complex system failure you managed.,How do you prioritize incident response?,How do you balance reliability with new features?,Explain your experience with reducing operational toil.
Frequently asked questions
- What are the primary responsibilities of a Site Reliability Engineer at MANTECH?
- As a Site Reliability Engineer at MANTECH, you will be responsible for ensuring the operational reliability, scalability, monitoring, and incident response for enterprise AI systems. This includes applying core reliability engineering principles, managing incident response and root cause analysis, implementing observability solutions, overseeing capacity planning and FinOps, and defining SLOs/SLAs.
- What technical skills are essential for the Site Reliability Engineer role at MANTECH?
- Essential technical skills for this Site Reliability Engineer role include extensive experience with cloud-native infrastructure, particularly Kubernetes. You'll also need deep knowledge of monitoring, alerting, and logging systems, and proven ability to automate operational tasks. Hands-on experience with the full observability stack (OpenTelemetry, Prometheus, Grafana, Loki, Tempo) and FinOps is preferred.
- Does this Site Reliability Engineer position require a security clearance?
- Yes, for onsite work, a TS/SCI clearance with Poly will be required for this Site Reliability Engineer position at MANTECH. Candidates must meet these clearance requirements to be considered for onsite duties.
- What is the educational and experience requirement for the Site Reliability Engineer role?
- The minimum qualifications for this Site Reliability Engineer role include a Bachelor’s degree in Computer Science, Engineering, or a related technical discipline, along with 5 or more years of experience in Site Reliability Engineering (SRE), DevOps, or production operations.
- How does MANTECH utilize Site Reliability Engineering principles in their AI initiatives?
- MANTECH utilizes Site Reliability Engineering principles to ensure the rapid design, deployment, operation, and sustainment of enterprise-scale AI, data, and mission platform capabilities. This involves focusing on operational outcomes, optimizing system performance, and maintaining high availability and stability of production systems supporting these AI efforts.