
Cloud Systems Architect (Remote)
Hire Feed · United States
- Hybrid
- Contract
- $120,000 / year
- United States
Job highlights
- Engineer AI systems infrastructure for reliability.
- Deploy and maintain Linux, Kubernetes, Prometheus.
- Automate processes, monitor system health.
- Collaborate with dev and ops teams.
- Contribute to cutting-edge AI model development.
About the role
Site Reliability Engineer (LInE)
We are hiring for one of our clients, seeking a Site Reliability Engineer (LInE) to work on a contractor basis. As a Site Reliability Engineer, you will apply your expertise to help train next-generation AI systems, shaping how models learn, reason, and perform through high-quality, real-world input. This role offers a unique opportunity to contribute to the development of frontier AI models, leveraging your domain knowledge to drive innovation in the AI industry.
Key Responsibilities
- Design, implement, and maintain scalable infrastructure using Linux, Kubernetes, and Prometheus, ensuring seamless deployments and high system availability.
- Monitor system health, analyze performance metrics, and proactively address bottlenecks or potential failures, minimizing manual intervention and increasing system reliability.
- Automate operational processes to minimize manual intervention and increase system reliability, and respond swiftly to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures.
- Collaborate closely with development and operations teams to deliver seamless deployments and high system availability, creating comprehensive documentation and clear runbooks for operational excellence.
- Respond to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures, ensuring high system availability and minimizing downtime.
Required Skills & Qualifications
- Proven experience designing, implementing, and maintaining scalable infrastructure using Linux, Kubernetes, and Prometheus, with a strong understanding of system health monitoring and performance metrics analysis.
- Strong understanding of automation tools and technologies, with experience in automating operational processes to minimize manual intervention and increase system reliability.
- Excellent problem-solving skills, with the ability to analyze complex system issues, identify root causes, and develop effective solutions.
- Strong communication and collaboration skills, with the ability to work closely with development and operations teams to deliver seamless deployments and high system availability.
- Experience with comprehensive documentation and clear runbooks for operational excellence, with a strong attention to detail and ability to create clear, concise documentation.
More About the Opportunity
This role offers a unique opportunity to work with a global leader in the AI industry, leveraging your domain knowledge to drive innovation and shape the development of next-generation AI systems. You will have the opportunity to work on a global scale, collaborating with top experts and contributing to the creation of cutting-edge AI models.
Equal Opportunity Employer
We hire based on skills and expertise. All qualified candidates are welcome regardless of background, experience, or prior employment history. Applications are reviewed solely on demonstrated technical ability and qualifications.
Key skills/competency
- Site Reliability Engineering
- Linux
- Kubernetes
- Prometheus
- System Monitoring
- Automation
- Incident Response
- Root Cause Analysis
- Infrastructure Design
- AI Systems
Skills & topics
- Site Reliability Engineer
- SRE
- Linux
- Kubernetes
- Prometheus
- System Monitoring
- Automation
- Incident Response
- Cloud Infrastructure
- AI Systems
- Remote
- Contractor
How to get hired
- Tailor your resume: Highlight your Linux, Kubernetes, Prometheus, and automation experience.
- Showcase problem-solving: Quantify your successes in incident response and system reliability.
- Demonstrate collaboration: Emphasize teamwork with dev and ops for seamless deployments.
- Prepare for technical questions: Be ready to discuss system design and performance metrics.
- Research the company: Understand their focus on AI and next-generation systems.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the primary focus of the Site Reliability Engineer role at Hire Feed?
- The primary focus of this Site Reliability Engineer role is to ensure the reliability and scalability of infrastructure for training next-generation AI systems, utilizing technologies like Linux, Kubernetes, and Prometheus.
- Is this a remote position for the Site Reliability Engineer role?
- Yes, this is a remote position, offering the flexibility to work from anywhere. The role is specifically listed as 'Remote (Work from Anywhere)'.
- What technical skills are essential for the Site Reliability Engineer position?
- Essential technical skills include proven experience with Linux, Kubernetes, and Prometheus for infrastructure design and maintenance, along with strong capabilities in system health monitoring, performance metrics analysis, and automation.
- What kind of AI systems will I be working with as a Site Reliability Engineer?
- You will be working with next-generation AI systems, contributing to how models learn, reason, and perform by providing high-quality, real-world input to train these frontier AI models.
- What is expected in terms of incident response for this Site Reliability Engineer role?
- You will be expected to respond swiftly to incidents, conduct thorough root cause analysis, and drive continuous improvements in incident response procedures to ensure high system availability and minimize downtime.
- Does Hire Feed consider experience history for Site Reliability Engineer applications?
- No, Hire Feed emphasizes skills and expertise. All qualified candidates are welcome regardless of background, experience, or prior employment history, with applications reviewed based on demonstrated technical ability and qualifications.