
Machine Learning Engineer
Sundayy · United States
- Hybrid
- Full-time
- $130,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Develop distributed vLLM infrastructure for LLMs.
- Optimize scalable AI inference systems.
- Work with Kubernetes, GoLang, and Python.
- Enhance performance and resource efficiency.
- Collaborate on cutting-edge AI technology.
About the role
About The Company
Red Hat is a global leader in open-source software solutions, dedicated to providing innovative and reliable technology that empowers enterprises to succeed in a digital-first world. With a strong commitment to open-source principles, Red Hat fosters a collaborative environment where developers, partners, and customers work together to create scalable, flexible, and secure software platforms. The company’s extensive portfolio includes cloud infrastructure, automation, middleware, and management solutions that are trusted by organizations worldwide to drive digital transformation and operational excellence.About The Role
We are seeking a highly skilled Machine Learning Engineer to join our AI Inference Engineering team. In this role, you will focus on developing and optimizing distributed vLLM infrastructure within the llm-d project, contributing to the advancement of scalable inference systems for large language models (LLMs). Your work will involve designing, developing, and testing innovative features that enhance the performance, stability, and resource efficiency of AI inference deployments. Collaborating closely with cross-functional teams, open-source communities, and senior engineers, you will help shape the future of AI deployment in enterprise environments. This position offers an exciting opportunity to work at the forefront of AI technology, leveraging cloud-native architectures, high-performance computing, and distributed systems to solve complex challenges in scalable inference.Qualifications
- Proficiency in Python and/or GoLang or similar programming languages
- Experience with cloud-native Kubernetes service mesh technologies such as Istio, Cilium, Envoy (WASM filters), and CNI
- Understanding of Layer7 networking, HTTP/2, gRPC, API gateways, and reverse proxies
- Knowledge of serving runtime technologies for hosting LLMs, including vLLM, SGLang, TensorRT-LLM
- Excellent written and verbal communication skills for effective technical collaboration
- Ability to work independently in a fast-paced and dynamic environment
- Preferred: Proficiency in C, C++, or Rust
- Preferred: Experience with Kubernetes ecosystem, custom APIs, operators, and Gateway API inference extension
- Preferred: Knowledge of high-performance networking protocols such as UCX, RoCE, InfiniBand, RDMA
- Preferred: Experience with GPU performance profiling tools like NVIDIA Nsight and distributed tracing techniques such as OpenTelemetry
- Strong understanding of computer architecture, parallel processing, and distributed computing concepts
- Educational background in computer science or related fields is advantageous, though hands-on experience is prioritized
- Active engagement in the ML research community through publications, conferences, or open-source contributions is a plus
Responsibilities
- Contribute to the design, development, and testing of new features and solutions for RedHat AI Inference platform
- Participate in upstream open-source communities to drive innovation in inference technologies
- Develop and maintain distributed inference infrastructure utilizing Kubernetes APIs, operators, and Gateway Inference Extension API for scalable LLM deployments
- Implement and maintain system components in Go and/or Rust to facilitate integration with vLLM and manage distributed inference workloads
- Design and optimize KV cache-aware routing and scoring algorithms to enhance memory utilization and request distribution at scale
- Enhance resource utilization, fault tolerance, and overall stability of inference systems
- Develop and evaluate various inference optimization algorithms to improve performance
- Engage in technical design discussions and share knowledge to foster continuous improvement
- Collaborate effectively with engineering and cross-functional teams to meet project objectives
- Participate in code reviews, provide constructive feedback, and ensure high-quality deliverables
- Seek mentorship from senior team members and contribute to a culture of learning and innovation
Benefits
- Comprehensive medical, dental, and vision insurance coverage
- Flexible Spending Account (FSA) for healthcare and dependent care expenses
- Health Savings Account (HSA) with high deductible medical plans
- Retirement savings plan with employer matching contributions
- Paid time off and holiday leave to promote work-life balance
- Paid parental leave for new parents
- Leave benefits including disability, family medical leave, and military leave
- Employee stock purchase plan and family planning reimbursement
- Tuition reimbursement, transportation expense accounts, and employee assistance programs
Equal Opportunity
Red Hat is committed to fostering an inclusive and diverse workplace. We are an equal Opportunity Employer and do not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, citizenship, age, veteran status, genetic information, physical or mental disability, medical condition, marital status, or any other protected characteristic. We support individuals with disabilities by providing reasonable accommodations throughout the application process. All applications are reviewed without bias, and we strive to create an environment where everyone can thrive and contribute to our mission of open-source innovation.Key skills/competency
Machine Learning Engineer, AI Inference, LLM, vLLM, Kubernetes, GoLang, Python, Distributed Systems, High-Performance Computing, Open SourceSkills & topics
- Machine Learning Engineer
- AI
- LLM
- vLLM
- Kubernetes
- GoLang
- Python
- Distributed Systems
- Inference Optimization
- Open Source
- Cloud Native
- Developer
How to get hired
- Research Red Hat's culture: Understand their commitment to open source and innovation.
- Tailor your resume: Highlight experience with Python, GoLang, Kubernetes, and LLM technologies.
- Showcase open-source contributions: Emphasize any involvement in ML research communities or projects.
- Prepare for technical interviews: Be ready to discuss distributed systems, networking, and inference optimization.
- Articulate your collaborative skills: Demonstrate your ability to work effectively with diverse teams.
Technical preparation
Master Python and GoLang for backend development.,Study Kubernetes for scalable infrastructure.,Understand LLM serving and inference techniques.,Practice distributed systems and networking concepts.
Behavioral questions
Describe a challenging ML project you led.,How do you handle complex technical problems?,Discuss a time you collaborated with others.,How do you stay updated with AI advancements?
Frequently asked questions
- What is the primary focus of the Machine Learning Engineer role at Red Hat?
- The Machine Learning Engineer at Red Hat will focus on developing and optimizing distributed vLLM infrastructure within the llm-d project, aiming to advance scalable inference systems for large language models (LLMs).
- What programming languages are essential for this Machine Learning Engineer position?
- Proficiency in Python and/or GoLang is essential. Experience with C, C++, or Rust is also preferred for this role.
- Does Red Hat require specific experience with Kubernetes for this Machine Learning Engineer role?
- Yes, experience with cloud-native Kubernetes service mesh technologies like Istio, Cilium, Envoy, and CNI is required. Experience with the broader Kubernetes ecosystem, custom APIs, operators, and Gateway API inference extension is preferred.
- What kind of LLM serving runtimes are relevant for this Machine Learning Engineer job?
- Knowledge of serving runtime technologies for hosting LLMs, including vLLM, SGLang, and TensorRT-LLM, is important for this Machine Learning Engineer position.
- Are contributions to open-source communities valued for this Machine Learning Engineer role?
- Yes, active engagement in the ML research community through publications, conferences, or open-source contributions is considered a plus for this Machine Learning Engineer role at Red Hat.
- What are the key responsibilities of the Machine Learning Engineer at Red Hat?
- Key responsibilities include designing, developing, and testing features for Red Hat's AI Inference platform, participating in open-source communities, maintaining distributed inference infrastructure, optimizing LLM deployments, and enhancing system stability and performance.
- What kind of networking knowledge is beneficial for the Machine Learning Engineer at Red Hat?
- Understanding of Layer7 networking, HTTP/2, gRPC, API gateways, and reverse proxies is required. Knowledge of high-performance networking protocols like UCX, RoCE, InfiniBand, and RDMA is also preferred.
- Does Red Hat offer remote work options for this Machine Learning Engineer position?
- The job description does not explicitly state remote work options, but Red Hat is a global company known for flexibility. It's best to clarify this during the application process.