
Senior Quality Assurance Engineer
Peach Pilot · United States
- Hybrid
- Contract
- $120,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Lead QA for AI systems and platform.
- Write test code and design evaluation pipelines.
- Focus on AI agent and knowledge graph quality.
- Collaborate directly with the founding engineering team.
- Shape product quality for executive clients.
About the role
About Peach Pilot
Peach Pilot builds a platform that ingests everything about how a company operates – every system, every process, every signal – and constructs a Company Brain: a living knowledge graph that connects people, decisions, and outcomes across the entire organization. We deploy 92 pre-built AI agents that work together across every business function, governed by humans at every critical step. The system gets smarter with every interaction. We don't sell software licenses. We embed into a client's operation, learn their business in weeks, show them what's broken backed by their own data, and redesign their highest-impact business functions with AI. Our first vertical is insurance. Our first client engagement is already scoped and funded. Peach Pilot is co-founded by Mario Montag (Predikto, acquired by a Fortune 50; McKinsey, PwC) and JP James (Hive Financial Assets, Georgia Tech, TITAN 100). We have a working platform with live infrastructure and a proven data-to-insights methodology.The Role
This is a hands-on contract QA role. You'll write test code, design evaluation pipelines, and set the quality bar before our platform reaches a client, working directly with the founding engineering team. We are not looking for someone who manages spreadsheets and delegates everything. We are looking for someone who can do the work and knows what good looks like. At Peach Pilot, quality is not just about whether buttons work. You are validating whether AI-generated analysis and agent recommendations are accurate enough to show to a CFO or CEO. One wrong finding in a Company X-Ray – the core deliverable that drives every client engagement – can break trust that took weeks to build. You are the last line of defense before our platform reaches a client's desk.The Challenge: QA for AI Is a Different Problem
Traditional QA assumes deterministic outputs. AI agents don't give you that. You will be validating quality in an environment where: 92 AI agents coordinate across business functions. Agent outputs must be accurate, auditable, and aligned with human-in-the-loop governance at every critical step. Multi-model routing (Claude, GPT, and others) means the same input can produce different outputs depending on which model handles it, and all of them need to meet the same quality bar. The Company X-Ray is our highest-stakes deliverable: a detailed analysis of a client's operations backed by their own data. Every finding must be reliable before it goes in front of a leadership team. Your end users are CEOs and operations leaders who have never used a terminal. A confusing output or a wrong recommendation doesn't just create a bug ticket, it kills adoption.What You'll Own & Test
Establish the Testing Foundation
- Establish the testing framework: unit, integration, end-to-end, and AI-specific evaluation pipelines using Playwright and Vitest.
- Define quality standards, test coverage requirements, and documentation practices in partnership with the Lead Engineer.
- Audit the existing platform and identify the highest-risk surfaces before the next client deployment.
AI Agent & Knowledge Graph Testing
- Design evaluation frameworks for non-deterministic LLM outputs — including prompt regression testing, model drift detection, and output quality scoring.
- Build automated test suites for the agent orchestration layer, including governance-agent audit-trail integrity and human-override behavior.
- Validate the Company Brain (Memgraph + Qdrant) for data accuracy, retrieval quality, and failure modes under real enterprise data, including entity resolution across systems and temporal data patterns.
- Test the Analysis Engine pipeline that surfaces Company X-Ray findings, ensuring insights are not just technically accurate but reliable enough to present to a client.
Platform & Integration Testing
- Own end-to-end testing of the data ingestion pipelines that connect to client systems (CRM, email, calls, calendars, documents, financial systems) through Nango's 700+ connector integration layer.
- Test multi-model routing logic to confirm cost-optimized task allocation behaves correctly across LLM providers via LiteLLM.
- Validate streaming response handling, latency thresholds, and graceful degradation when a model is unavailable or slow.
- Own file ingestion pipeline testing (Word, Excel, PowerPoint, PDF) including encryption, formatting edge cases, and audit-trail continuity.
Who You Are
- 7+ years of QA engineering experience, with at least 3 years in a senior or lead capacity where you shaped process and standards, not just executed them.
- You have tested AI/LLM-powered applications. You understand prompt sensitivity, output variance, and how to build eval pipelines that catch regressions across model updates.
- You speak in ownership: you've built the eval pipeline, owned model quality, gated the release — not just run someone else's test suite.
- You write test code. Python is your primary tool. You have built and maintained CI/CD-integrated test suites, and you don't wait for someone to file a bug to find one.
- Hands-on experience with Playwright and Vitest in a production environment, and you've built automation frameworks from scratch, not just inherited them.
- Comfortable testing complex API chains, async/streaming responses, and multi-service workflows. Data pipelines and knowledge graph outputs don't intimidate you.
- You test for confusion and trust failure, not just broken functionality. Your end users are non-technical executives, and you advocate for them.
- US-based, able to overlap roughly 5 hours per day with EDT, and available for full-time contract hours.
The Stack You'll Test Against
- AI/LLM: Anthropic Claude, OpenAI GPT, LiteLLM (multi-model routing), custom agent orchestration with reinforcement learning
- Backend: Python (FastAPI), async agent runtime, Pydantic
- Data & Graph: Memgraph · Neo4j · Qdrant · PostgreSQL · Redis
- Frontend: React/Next.js, TypeScript, Tailwind CSS
- Integrations: Nango (700+ connectors)
- Infrastructure: Google Cloud Platform (Cloud Run, GCE, Firebase) · GitHub Actions CI/CD · Docker. We run on GCP
- Testing: Playwright, Vitest
Even Better If
- You have experience with LLM evaluation frameworks (e.g., LangSmith, DeepEval, Promptfoo, RAGAS, or custom eval pipelines).
- You have tested agent frameworks or orchestration layers in a production environment.
- You have a background in a regulated industry (insurance, finance, healthcare) where audit-trail integrity is non-negotiable.
- You have worked alongside Forward Deployed or solutions engineering teams and understand field deployment risk.
What Makes This Different
You'll work directly with the founding engineering team on a platform that already has live infrastructure and a first client engagement in motion. Your findings shape the product before it reaches the people who matter most — the executives who decide whether to trust what our agents put in front of them. Every engagement makes the platform smarter, and you are the reason a client can act on what they see.Compensation & Engagement
- Compensation: 1099 hourly rate depending on experience (contract).
- Structure: Contract, full-time hours (~40/week). Contract-to-hire possible if the engagement goes well.
- Location: Fully Remote US Based, with ~5 hours/day of EDT overlap (Atlanta, GA (Buckhead) candidates welcome).
Key skills/competency
- Senior QA Engineer
- AI Systems Testing
- LLM Evaluation
- Python
- Playwright
- Vitest
- API Testing
- Data Pipelines
- Knowledge Graphs
- CI/CD
Skills & topics
- Senior QA Engineer
- AI
- LLM
- Python
- Playwright
- Vitest
- Quality Assurance
- Software Testing
- Evaluation Frameworks
- Remote
How to get hired
- Tailor your resume: Highlight AI/LLM testing, Python, Playwright, and Vitest experience.
- Showcase ownership: Emphasize building eval pipelines and gating releases.
- Prepare for technical interviews: Be ready to discuss AI quality challenges and solutions.
- Demonstrate executive advocacy: Explain how you test for non-technical end-users.
- Highlight remote work skills: Showcase your ability to collaborate effectively remotely.
Technical preparation
Master Python for test automation.,Deepen Playwright and Vitest knowledge.,Understand LLM evaluation techniques.,Practice testing complex API chains.
Behavioral questions
Describe owning a testing process end-to-end.,How do you ensure quality for AI outputs?,How do you advocate for non-technical users?,Share an experience building a test framework.
Frequently asked questions
- What is Peach Pilot's primary focus as a company?
- Peach Pilot focuses on transforming how businesses operate by building a 'Company Brain' platform that uses AI agents to analyze and improve business functions, starting with the insurance industry.
- What makes QA for AI systems different at Peach Pilot?
- QA for AI at Peach Pilot involves testing non-deterministic outputs from AI agents, ensuring accuracy and auditability across multiple LLM models, and validating the reliability of insights for executive decision-making, unlike traditional deterministic QA.
- What are the core responsibilities of the Senior QA Engineer?
- The Senior QA Engineer will establish the testing foundation, design evaluation frameworks for AI outputs, test AI agents and knowledge graphs, and perform end-to-end testing of data ingestion and platform integrations.
- What programming languages and tools are essential for this role?
- Python is the primary tool for writing test code. Experience with Playwright and Vitest for building automation frameworks and CI/CD integration is crucial.
- What kind of experience is Peach Pilot looking for in a Senior QA Engineer?
- Peach Pilot seeks 7+ years of QA experience, with at least 3 years in a senior/lead capacity, proven experience testing AI/LLM applications, and a strong ability to write test code, particularly in Python.
- Can this contract role lead to a full-time position?
- Yes, a contract-to-hire possibility exists if the engagement goes well, offering a potential path to a full-time role within the company.
- What is the expected time commitment and location for this contract?
- This is a full-time contract role (~40 hours/week) for a US-based remote candidate, requiring approximately 5 hours of overlap per day with EDT.
- What does 'human-in-the-loop governance' mean in the context of Peach Pilot's platform?
- Human-in-the-loop governance means that human oversight and decision-making are integrated at critical steps of the AI agent's processes to ensure accuracy, alignment, and trust before AI-driven outcomes are finalized.
- How does Peach Pilot ensure the reliability of the 'Company X-Ray' deliverable?
- Reliability is ensured through rigorous testing of the Analysis Engine pipeline, validating data accuracy and retrieval quality, and confirming that AI-generated findings are dependable enough for presentation to client leadership.
- What are the advantages of working on the early engineering team at Peach Pilot?
- Working on the early engineering team provides direct collaboration with founders, the opportunity to shape product quality from the ground up, and the chance to influence a platform with live infrastructure and an active client engagement.
Similar roles
Open positions we recommend based on this role.