PitchMeAI
MDCalc

QA Engineer, AI Products

MDCalc · United States

  • Hybrid
  • Full-time
  • $130,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Ensure quality of AI-powered medical features.
  • Test LLM systems with non-deterministic outputs.
  • Develop automated evaluation pipelines.
  • Collaborate with engineers, product, and clinical teams.
  • Define and monitor AI quality metrics.

About the role

The Opportunity

Since 2005, MDCalc has been an essential part of the clinician’s workflow to help achieve better patient outcomes. Actively used by more than 65% of physicians worldwide, MDCalc is the most broadly used medical reference – at the point-of-care – for clinical decision tools and content, and one of only four references used by >50% of US HCPs. These evidence-based tools and content are used by millions of medical professionals globally and support 50+ specialties and cover 200+ patient conditions.

To continue to further accelerate and steward this growth, we are expanding the AI product team with a QA Engineer. This role will be critical to MDCalc’s expanded success in continuing to support our millions of clinical users worldwide in taking care of hundreds of millions of patients.

The Role

As a QA Engineer on the AI Products group at MDCalc, you will play a key role in ensuring the quality, reliability, and clinical trustworthiness of MDCalc's AI-powered features. You'll focus on the unique challenges of testing LLM-based systems, where outputs are non-deterministic, correctness is often a spectrum rather than a binary, and regressions can be subtle. You'll be part of a collaborative, fast-moving team that takes pride in delivering software that clinicians trust to care for millions of patients worldwide.

Responsibilities:

  • Design and execute test strategies for LLM-powered features, including prompt regression testing, output evaluation, and hallucination detection
  • Build and maintain automated evaluation pipelines (eval sets, golden datasets, LLM-as-judge frameworks) to catch quality regressions in non-deterministic outputs
  • Perform black-box and exploratory testing of MDCalc's AI features across web and mobile, with particular attention to clinical accuracy, safety, and edge cases
  • Define quality metrics for AI outputs (accuracy, faithfulness, relevance, safety, latency, cost) and establish thresholds for release readiness
  • Collaborate cross-functionally with engineers, product managers, ML/AI engineers, and clinical reviewers to define what "good" looks like for AI responses
  • Investigate and triage AI failure modes, distinguishing model issues, prompt issues, retrieval issues, and integration bugs
  • Participate in team discussions, offering feedback on testability, risks, prompt design, and guardrails
  • Help develop QA strategies to expand future testing capacity, automation, and evaluation coverage as the AI product surface grows

Your Background

  • 5+ years of experience in software QA, with at least 1 year of hands-on testing of LLM-based or AI/ML-powered features
  • Strong understanding of QA principles, test case creation/documentation, and best practices for both deterministic and non-deterministic systems
  • Hands-on experience with LLM tooling and concepts: prompt engineering, RAG systems, evaluation frameworks (e.g., Promptfoo, Braintrust, LangSmith, DeepEval, Ragas, OpenAI Evals), and LLM APIs (OpenAI, Anthropic, etc.)
  • Experience designing automated qualitative evaluation approaches, including LLM-as-judge, rubric-based scoring, semantic similarity checks, and golden dataset regression testing
  • Proficiency with test automation tools, with a focus on Playwright
  • Strong SQL skills for data validation, test data creation, and verifying data integrity across systems
  • Familiarity with token usage, latency profiling, and cost monitoring as quality signals
  • Eagerness to learn quickly and a positive, solutions-oriented attitude
  • Clear and concise communicator, able to surface issues, blockers, and risks effectively when communicating ambiguous or probabilistic failures
  • Self-motivated, proactive, and able to manage time and priorities independently

What MDCalc Offers

  • Ability to make a true difference in medicine: MDCalc is the most broadly used medical reference by physicians, used by over 65% of US attending doctors weekly
  • Medical, Dental, & Vision Coverage, with option to extend to your dependents
  • Company-sponsored short-term insurance
  • Fully-paid 8 week parental leave, after 6 months of employment
  • Company-sponsored 401k, after 3 months of employment
  • Unlimited vacation for salaried roles - we trust you to take the time you need
  • Bi-annual company offsites to connect, reflect, and plan together
  • Work from home monthly stipend
  • A culture of fun and motivated team members who believe in a greater mission here at MDCalc

Key skills/competency

  • AI Testing
  • LLM Evaluation
  • Prompt Engineering
  • QA Automation
  • Playwright
  • SQL
  • Software Quality Assurance
  • Test Strategy
  • Regression Testing
  • Clinical Accuracy

Skills & topics

  • QA Engineer
  • AI Products
  • Machine Learning
  • LLM
  • Prompt Engineering
  • Software Quality Assurance
  • Test Automation
  • Playwright
  • SQL
  • Healthcare Technology

How to get hired

  • Tailor your resume: Highlight AI/LLM testing experience, Playwright, SQL, and prompt engineering skills.
  • Quantify achievements: Use numbers to showcase your impact on quality and efficiency.
  • Showcase understanding: Emphasize your knowledge of LLM concepts and evaluation frameworks.
  • Prepare for technical questions: Be ready to discuss AI failure modes and testing strategies.
  • Demonstrate cultural fit: Highlight your collaborative spirit and solutions-oriented attitude.

Technical preparation

Practice prompt engineering techniques.,Build test cases for LLM outputs.,Automate tests using Playwright.,Query databases with complex SQL.

Behavioral questions

Describe a complex bug you found.,How do you handle ambiguous requirements?,Share an experience collaborating with engineers.,How do you prioritize testing efforts?

Frequently asked questions

What are the key responsibilities for a QA Engineer, AI Products at MDCalc?
The QA Engineer will design and execute test strategies for LLM-powered features, build automated evaluation pipelines, perform black-box and exploratory testing, define quality metrics, and collaborate with cross-functional teams to ensure the reliability and clinical trustworthiness of MDCalc's AI features.
What specific experience is required for the QA Engineer, AI Products role at MDCalc?
We require 5+ years of software QA experience, with at least 1 year focused on testing LLM-based or AI/ML-powered features. Hands-on experience with LLM tooling, prompt engineering, evaluation frameworks, LLM APIs, Playwright, and strong SQL skills are essential.
How does MDCalc ensure the clinical accuracy of its AI features?
MDCalc emphasizes clinical accuracy and safety through rigorous testing, including evaluating AI outputs for correctness, hallucination detection, and edge cases. The role involves close collaboration with clinical reviewers and defining specific quality metrics for AI responses.
What kind of team will I be joining as a QA Engineer at MDCalc?
You will join a collaborative, fast-moving AI Products team that takes pride in delivering high-quality software. The team is composed of engineers, product managers, ML/AI engineers, and clinical reviewers working towards a shared mission.
What are the benefits of working at MDCalc as a QA Engineer?
MDCalc offers the chance to make a significant impact in medicine, comprehensive medical, dental, and vision coverage, paid parental leave, 401k sponsorship, unlimited vacation, and a supportive culture with fun and motivated team members.
Is this a remote or hybrid position for the QA Engineer, AI Products role?
The job description mentions a 'Work from home monthly stipend,' suggesting a strong remote or hybrid work arrangement is supported, though the exact nature should be clarified during the application process.
What specific LLM evaluation frameworks are preferred for this QA Engineer position?
Experience with evaluation frameworks such as Promptfoo, Braintrust, LangSmith, DeepEval, Ragas, and OpenAI Evals is highly valued for this role.
How important is SQL for the QA Engineer, AI Products role at MDCalc?
Strong SQL skills are crucial for data validation, creating test data, and verifying data integrity across systems, making it a key requirement for this position.