
Site Reliability Engineer
micro1 · United States
- Hybrid
- Contract
- $140,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Optimize AI model training in containerized environments.
- Resolve complex infrastructure issues rapidly.
- Implement resilient system recovery protocols.
- Automate workflows and CI/CD pipelines.
- Manage critical, long-running server processes.
About the role
DevOps Engineer
Join our customer's team as a DevOps Engineer on a specialized, high-intensity project dedicated to training and optimizing AI models within advanced containerized environments. This is an expert-level, terminal-intensive engagement where your ability to troubleshoot, recover, and optimize dynamic infrastructure will directly influence project success. Demonstrate elite technical execution in a role that offers the potential for further engagement or transition to subsequent project phases.
Key Responsibilities
- Architect, maintain, and optimize containerized environments for large-scale AI model training and data processing.
- Rapidly diagnose and resolve issues in live systems, employing advanced terminal-native problem-solving skills.
- Implement dynamic infrastructure recovery protocols to ensure high system availability and resilience.
- Collaborate closely with cross-functional teams to streamline CI/CD pipelines and automate critical workflows.
- Manage, monitor, and troubleshoot long-running server processes, proactively identifying and addressing resource bottlenecks and failures.
- Replan and recover interrupted processes in Dockerized sandboxes, minimizing downtime and maximizing efficiency.
- Contribute technical expertise to system builds, server administration, and infrastructure management for cutting-edge AI workloads.
Required Skills and Qualifications
- Proven expertise as a DevOps Engineer in production environments, with hands-on terminal proficiency.
- Mastery of dynamic infrastructure recovery, error resolution, and live process management in complex, containerized setups.
- Demonstrated skill in Docker, Kubernetes, and related container orchestration technologies.
- Strong programming abilities in Python, with proficiency in Bash scripting and familiarity with JavaScript/TypeScript, Go, Rust, or C/C++.
- Experience with build systems, package managers, databases, web servers, ML frameworks, version control, and cryptography tools.
- Exceptional troubleshooting ability, especially in multi-step, mid-execution replanning scenarios.
- Systems-first mindset with a passion for optimizing large-scale, mission-critical environments.
Preferred Qualifications
- Hands-on experience supporting AI/ML model training pipelines or high-availability compute clusters.
- Background in security, cryptography, or compliance within containerized or cloud-driven environments.
- Contributions to open-source DevOps tooling or infrastructure projects.
Key skills/competency
- DevOps Engineer
- Site Reliability Engineering
- AI Model Training
- Containerization
- Docker
- Kubernetes
- Python
- Bash Scripting
- System Administration
- Infrastructure Optimization
Skills & topics
- DevOps Engineer
- Site Reliability Engineer
- AI
- Machine Learning
- Containerization
- Docker
- Kubernetes
- Python
- Bash
- Cloud Computing
- Infrastructure
- System Administration
- Remote
- Contractor
How to get hired
- Tailor your resume: Highlight DevOps experience and AI/ML infrastructure skills.
- Showcase terminal proficiency: Emphasize direct experience with command-line tools.
- Demonstrate container expertise: Detail your work with Docker and Kubernetes.
- Prepare for technical interviews: Be ready to discuss complex troubleshooting scenarios.
- Highlight Python/Bash skills: Provide examples of automation and scripting projects.
Technical preparation
Master Docker and Kubernetes operations.,Practice complex Bash scripting scenarios.,Deepen Python knowledge for automation.,Study AI/ML infrastructure challenges.
Behavioral questions
Describe a critical system recovery.,How do you troubleshoot complex issues?,Share an experience optimizing infrastructure.,How do you collaborate with other teams?
Frequently asked questions
- What specific AI/ML workloads will I be supporting as a DevOps Engineer at micro1?
- As a DevOps Engineer at micro1, you will focus on training and optimizing AI models. This involves working within advanced containerized environments to manage large-scale data processing and model development pipelines.
- Is this a remote DevOps Engineer role with micro1?
- Yes, this DevOps Engineer position at micro1 is fully remote, allowing you to work from anywhere.
- What level of terminal proficiency is expected for this DevOps Engineer role?
- Exceptional, hands-on terminal proficiency is a core requirement. You will be expected to diagnose and resolve issues directly via the command line in complex, dynamic systems.
- What are the key container technologies used for this DevOps Engineer position?
- The primary container technologies you will work with are Docker and Kubernetes, along with related orchestration tools, to manage large-scale AI training environments.
- Does micro1 offer opportunities for contract extension or future projects for DevOps Engineers?
- This engagement offers the potential for further contract extension or transition into subsequent project phases, providing ongoing opportunities for expert-level contributions.
- What programming languages are most important for this DevOps Engineer role?
- Strong Python programming abilities and proficiency in Bash scripting are essential. Familiarity with JavaScript/TypeScript, Go, Rust, or C/C++ is also beneficial.
- What kind of troubleshooting experience is crucial for this role?
- Exceptional troubleshooting skills are critical, particularly in multi-step, mid-execution replanning scenarios within dynamic, containerized systems.