PitchMeAI
Q6 Cyber

SR. CLOUD INFRASTRUCTURE ENGINEER (AI & LLM PLATFORMS)

Q6 Cyber · United States

  • Hybrid
  • Full-time
  • $150,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Build AI infrastructure for LLM integration.
  • Manage RAG pipelines and MCP servers.
  • Deploy and scale using Kubernetes and Terraform.
  • Ensure security for LLM data workflows.
  • Collaborate with cross-functional teams.

About the role

Cloud Infrastructure Engineer AI LLM Platforms

We are seeking a specialized Infrastructure Engineer to bridge the gap between our large data repositories, Cloud Platform and the rapidly evolving world of Large Language Models (LLMs). You will be responsible for building the "plumbing" that allows our internal teams and external users to leverage AI effectively. This includes deploying Model Context Protocol (MCP) servers, building agentic execution environments, and scaling our internal Retrieval-Augmented Generation (RAG) architecture.

Roles and Responsibilities

Key Responsibilities

  • AI Architecture Guidance: Guide the architecture that will allow us to leverage AI tools with our large existing data stores and incoming streams of realtime intelligence.
  • Cross-Team Integration: Work closely with other infrastructure engineers and software development teams to integrate AI tools into existing systems.
  • MCP Ecosystem Management: Design, deploy, and maintain Model Context Protocol (MCP) servers to allow LLMs to securely interact with our internal databases, APIs, and external tooling.
  • Agentic Infrastructure: Build and orchestrate sandboxed, scalable environments (e.g., using Docker or specialized runtimes) where users can safely build and execute AI agents.
  • Internal RAG Platform: Develop and manage the infrastructure for our internal RAG (Retrieval-Augmented Generation) pipeline, including vector database management (e.g., Pinecone, Weaviate, or pgvector) and automated embedding pipelines.
  • Deployment & Scaling: Utilize Kubernetes (K8s) and Infrastructure as Code (Terraform/Pulumi) to deploy LLM-related tools, ensuring high availability and low latency for model inference and data retrieval.
  • Security & Governance: Implement strict guardrails for data privacy within LLM workflows, ensuring internal datasets remain secure while being accessible to authorized AI tools.

Required Qualifications

  • 5+ years of experience in DevOps, Platform Engineering, or SRE, with at least 1-2 years specifically focused on AI/ML infrastructure.
  • Proven track record of building production-grade RAG pipelines or LLM-integrated applications.
  • Thrives in "day zero" environments where the tools and protocols (like MCP) are evolving weekly.
  • Deep understanding of the security implications of LLMs (prompt injection, data leakage, and secure tool execution).
  • Experience working with substantial datasets (over 1bn objects, dozens or hundreds of TBs) and the challenges of leveraging AI tools with these data sets.
  • Bachelor's degree or equivalent in computer science or related field.

Required Technical Skills

  • Cloud & Orchestration: AWS/GCP/Azure, Kubernetes, Terraform, Helm.
  • AI Frameworks: LangChain, LlamaIndex, LangGraph.
  • Data & Vectors: Pinecone, Milvus, Qdrant, or pgvector; Apache Kafka/Pulsar; Elasticsearch/OpenSearch; traditional SQL RDBMS.
  • Languages: Python (Expert), TypeScript/Node.js (for MCP development), Go.
  • AI Protocols: Model Context Protocol (MCP), REST/gRPC.

Key skills/competency

  • Cloud Infrastructure Engineering
  • AI/ML Infrastructure
  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • Kubernetes
  • Terraform
  • Python
  • Data Engineering
  • Security
  • DevOps

Skills & topics

  • Cloud Infrastructure Engineer
  • AI
  • LLM
  • DevOps
  • SRE
  • Platform Engineering
  • Kubernetes
  • Terraform
  • Python
  • RAG

How to get hired

  • Tailor your resume: Highlight your experience with AI/ML infrastructure, RAG, and LLMs. Quantify achievements related to large datasets and production deployments.
  • Showcase technical skills: Emphasize proficiency in Kubernetes, Terraform, Python, and specific AI frameworks like LangChain.
  • Demonstrate LLM security understanding: Detail your knowledge of prompt injection, data leakage, and secure tool execution.
  • Highlight "day zero" experience: Mention your comfort and success in rapidly evolving environments with new protocols.
  • Prepare for technical interviews: Be ready to discuss your experience with large-scale data, cloud platforms, and infrastructure as code.

Technical preparation

Master Python for AI/ML development.,Deep dive into Kubernetes and Terraform.,Practice with vector databases like Pinecone.,Understand RAG and MCP protocol concepts.

Behavioral questions

Describe a challenging 'day zero' environment.,How do you ensure LLM data security?,Explain integrating AI into existing systems.,How would you scale RAG infrastructure?

Frequently asked questions

What is the primary focus of the Cloud Infrastructure Engineer AI LLM Platforms role at Q6 Cyber?
The primary focus is on building and managing the infrastructure required to integrate Large Language Models (LLMs) with Q6 Cyber's extensive data repositories and real-time intelligence streams, enabling effective AI utilization.
What specific AI/ML infrastructure experience is required for this role at Q6 Cyber?
A minimum of 1-2 years specifically focused on AI/ML infrastructure is required, in addition to 5+ years in broader DevOps, Platform Engineering, or SRE roles. Experience with RAG pipelines or LLM-integrated applications is crucial.
How does Q6 Cyber approach the rapidly evolving nature of AI tools and protocols like MCP?
Q6 Cyber seeks individuals who thrive in 'day zero' environments where tools and protocols are constantly evolving, indicating a culture that embraces rapid change and innovation in the AI space.
What are the key technical skills needed for the Cloud Infrastructure Engineer AI LLM Platforms position?
Key technical skills include cloud platforms (AWS/GCP/Azure), Kubernetes, Terraform, Helm, AI frameworks (LangChain, LlamaIndex), vector databases (Pinecone, Weaviate), and programming languages like Python and Go.
How important is data security in the LLM workflows at Q6 Cyber?
Data security and governance are critical. The role involves implementing strict guardrails for data privacy within LLM workflows to ensure internal datasets remain secure while accessible to authorized AI tools.
What kind of data scale can I expect to work with as a Cloud Infrastructure Engineer at Q6 Cyber?
You can expect to work with substantial datasets, often exceeding 1 billion objects and ranging from dozens to hundreds of terabytes, presenting unique challenges for AI tool integration.
What are the responsibilities related to agentic infrastructure at Q6 Cyber?
You will be responsible for building and orchestrating sandboxed, scalable environments for users to safely develop and execute AI agents, likely using containerization technologies like Docker.