
SR. CLOUD INFRASTRUCTURE ENGINEER (AI & LLM PLATFORMS)
Q6 Cyber · United States
- Hybrid
- Full-time
- $150,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Build AI infrastructure for large datasets.
- Deploy MCP servers and agentic environments.
- Develop internal RAG pipeline infrastructure.
- Scale LLM tools using Kubernetes and Terraform.
- Ensure AI workflow security and data privacy.
About the role
About the Role
We are seeking a specialized Infrastructure Engineer to bridge the gap between our large data repositories, Cloud Platform and the rapidly evolving world of Large Language Models (LLMs). You will be responsible for building the "plumbing" that allows our internal teams and external users to leverage AI effectively. This includes deploying Model Context Protocol (MCP) servers, building agentic execution environments, and scaling our internal Retrieval-Augmented Generation (RAG) architecture.
Key Responsibilities
- AI Architecture Guidance: Guide the architecture that will allow us to leverage AI tools with our large existing data stores and incoming streams of realtime intelligence.
- Cross-Team Integration: Work closely with other infrastructure engineers and software development teams to integrate AI tools into existing systems.
- MCP Ecosystem Management: Design, deploy, and maintain Model Context Protocol (MCP) servers to allow LLMs to securely interact with our internal databases, APIs, and external tooling.
- Agentic Infrastructure: Build and orchestrate sandboxed, scalable environments (e.g., using Docker or specialized runtimes) where users can safely build and execute AI agents.
- Internal RAG Platform: Develop and manage the infrastructure for our internal RAG (Retrieval-Augmented Generation) pipeline, including vector database management (e.g., Pinecone, Weaviate, or pgvector) and automated embedding pipelines.
- Deployment & Scaling: Utilize Kubernetes (K8s) and Infrastructure as Code (Terraform/Pulumi) to deploy LLM-related tools, ensuring high availability and low latency for model inference and data retrieval.
- Security & Governance: Implement strict guardrails for data privacy within LLM workflows, ensuring internal datasets remain secure while being accessible to authorized AI tools.
Required Qualifications
- 5+ years of experience in DevOps, Platform Engineering, or SRE, with at least 1-2 years specifically focused on AI/ML infrastructure.
- Proven track record of building production-grade RAG pipelines or LLM-integrated applications.
- Thrives in "day zero" environments where the tools and protocols (like MCP) are evolving weekly.
- Deep understanding of the security implications of LLMs (prompt injection, data leakage, and secure tool execution).
- Experience working with substantial datasets (over 1bn objects, dozens or hundreds of TBs) and the challenges of leveraging AI tools with these data sets.
- Bachelor's degree or equivalent in computer science or related field.
Required Technical Skills
- Cloud & Orchestration: AWS/GCP/Azure, Kubernetes, Terraform, Helm.
- AI Frameworks: LangChain, LlamaIndex, LangGraph.
- Data & Vectors: Pinecone, Milvus, Qdrant, or pgvector; Apache Kafka/Pulsar; Elasticsearch/OpenSearch; traditional SQL RDBMS.
- Languages: Python (Expert), TypeScript/Node.js (for MCP development), Go.
- AI Protocols: Model Context Protocol (MCP), REST/gRPC.
Key skills/competency
- Cloud Infrastructure Engineering
- AI/ML Infrastructure
- Large Language Models (LLMs)
- Retrieval-Augmented Generation (RAG)
- Kubernetes
- Terraform
- Vector Databases
- Python
- DevOps
- Security & Governance
Skills & topics
- Cloud Infrastructure Engineer
- AI
- LLM
- RAG
- Kubernetes
- Terraform
- DevOps
- SRE
- Platform Engineering
- Python
- AWS
- GCP
- Azure
- Vector Databases
- MCP
How to get hired
- Tailor your resume: Highlight AI/ML infrastructure, RAG, and LLM experience. Emphasize production-grade deployments.
- Showcase technical skills: Detail your expertise in Kubernetes, Terraform, Python, and vector databases.
- Demonstrate problem-solving: Provide examples of working in evolving environments and managing large datasets.
- Prepare for technical interviews: Be ready to discuss AI architecture, security implications, and infrastructure scaling.
- Understand Q6 Cyber's mission: Research their focus on real-time intelligence and AI integration.
Technical preparation
Master Python for AI/ML development.,Practice Kubernetes and Terraform for deployment.,Build RAG pipelines with vector databases.,Study LLM security and MCP protocols.
Behavioral questions
Describe a challenging AI infrastructure project.,How do you handle rapidly evolving tech?,Explain your approach to AI data security.,How do you collaborate with software teams?
Frequently asked questions
- What is the primary focus of the Senior Cloud Infrastructure Engineer role at Q6 Cyber?
- The Senior Cloud Infrastructure Engineer role at Q6 Cyber focuses on building and scaling the infrastructure that enables the effective use of AI and Large Language Models (LLMs) with the company's extensive data repositories. This includes deploying MCP servers, creating agentic execution environments, and managing the internal RAG architecture.
- What specific AI technologies are central to this Senior Cloud Infrastructure Engineer position?
- Central AI technologies for this role include Retrieval-Augmented Generation (RAG) pipelines, Model Context Protocol (MCP) servers, and agentic execution environments. Familiarity with AI frameworks like LangChain, LlamaIndex, and LangGraph is also crucial.
- What kind of data infrastructure experience is Q6 Cyber looking for in this role?
- Q6 Cyber is looking for experience with substantial datasets (over 1 billion objects, dozens or hundreds of TBs) and the challenges associated with leveraging AI tools with such data. This includes expertise in vector databases (Pinecone, Weaviate, pgvector), real-time data streams (Kafka/Pulsar), and search/analytics engines (Elasticsearch/OpenSearch).
- How important is experience with containerization and Infrastructure as Code for this role?
- Containerization (Docker) and Infrastructure as Code (Terraform/Pulumi) are critically important. You will use Kubernetes for deployment and scaling of LLM-related tools, ensuring high availability and low latency. Terraform is expected for managing cloud infrastructure.
- What are the security considerations for an AI/LLM Infrastructure Engineer at Q6 Cyber?
- Security is paramount. You will implement strict guardrails for data privacy within LLM workflows, focusing on preventing prompt injection, data leakage, and ensuring secure tool execution for internal datasets while maintaining accessibility for authorized AI tools.
- Does Q6 Cyber expect this Senior Cloud Infrastructure Engineer to work with evolving technologies?
- Yes, the job description explicitly states that the ideal candidate thrives in "day zero" environments where tools and protocols are evolving weekly. This role requires adaptability and a proactive approach to learning and implementing new technologies in the AI space.
- What programming languages are essential for the Senior Cloud Infrastructure Engineer role?