Senior AI/ML Test and Evaluation Engineer
OpenTeams · Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO - Hybrid/Remote
Posted 14 days ago · $145,000–$250,000
or apply directly on OpenTeams's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 14 days OpenTeams's roles stay open a median of 14 days
- Reposted
- No
- Salary listed
- Yes 83% of OpenTeams's roles list one
- Ghost-job risk at OpenTeams
- moderate 2 stale, 1 reposted of 12 open
- Hiring momentum
- 15 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-09-17
Measured from postings appearing on and disappearing from OpenTeams's own greenhouse board since 2026-08-03. Full hiring picture for OpenTeams.
About this role
The Senior AI/ML Test and Evaluation Engineer will focus on building and operating benchmarking and evaluation capabilities for AI models. This role involves designing evaluation methodologies, creating automated metrics, and documenting limitations and failure modes of models. The engineer will also produce evaluation reports for senior stakeholders and support integration with partner organizations.
- benefits
- 4/5
- freshness
- 4/5
- career value
- 4/5
- role clarity
- 5/5
- pay transparency
- 5/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of OpenTeams as an employer.
What you need
- U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance
- 6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments
- Strong Python proficiency in a machine learning or data science context
- Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem
- Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks
- Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results
Nice to have
- Active U.S. security clearance
- Prior AI/ML evaluation or test and evaluation experience supporting the Department of Defense, Intelligence Community, or another federal customer
- Experience designing human-in-the-loop evaluations, measuring inter-rater reliability, or facilitating structured expert adjudication
- Experience defining or implementing benchmark interchange formats or evaluation standards used across multiple organizations
- Familiarity with intelligence analysis workflows or other high-stakes analytical domains
What you get
- Medical, Dental & Vision – 100% paid for employees, 75% for dependents
- 401(k) Match – Up to 5% with full vesting after 2 years
- Unlimited PTO – With a required minimum of 15 days off annually
- Fully Remote Setup – Includes up to $3,000 equipment reimbursement
- Continuous Education – Includes up to $500 reimbursement
- Disability & Life Insurance – 100% employer-paid
Worth weighing
- Role is contingent upon contract award
- Travel of up to 15% may be required, primarily to Government facilities and between company locations
- Focus on evaluating what models get wrong rather than what they get right may not appeal to all candidates
- Hands-on engineering with an open-source toolchain may require adaptability to evolving technologies
Summarised from OpenTeams's posting. Read the full original.
Listed by OpenTeams on their greenhouse job board, last confirmed open on 2026-09-17. PitchMeAI is not the employer.
More roles at OpenTeams
- Distributed Systems ML Infrastructure EngineerWashington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote
- Senior Engineering ArchitectUnited States - Remote
- Senior Infrastructure Engineer - AI/ML PlatformUnited States - Remote
- Site Reliability Engineer / DevSecOps EngineerWashington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote
- Program ManagerWashington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote
- Technical Delivery LeadWashington DC Metro OR Denver Metro – Hybrid
- Full-Stack Platform EngineerWashington, DC Metro; Denver, CO Metro; or Colorado Springs, CO - Hybrid/Remote
- Technical Project Manager - Project SuccessRemote - U.S