or apply directly on Phizenix's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 26 days Phizenix's roles stay open a median of 45 days
- Reposted
- No
- Salary listed
- No 0% of Phizenix's roles list one
- Ghost-job risk at Phizenix
- high 7 stale, 0 reposted of 13 open
- Hiring momentum
- 31 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-09-17
Measured from postings appearing on and disappearing from Phizenix's own greenhouse board since 2026-08-03. Full hiring picture for Phizenix.
About this role
The LLM / Agentic Evaluation Rig Engineer will be responsible for building and maintaining the evaluation infrastructure that ensures AI outputs meet quality standards before being shipped. This includes curating datasets, developing scoring systems for various quality metrics, and integrating evaluations into continuous integration processes. The role emphasizes rigorous measurement of grounding, faithfulness, and hallucination in AI outputs, alongside collaboration with other teams to enhance model performance.
- benefits
- 1/5
- freshness
- 4/5
- career value
- 4/5
- role clarity
- 5/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Phizenix as an employer.
What you need
- 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling.
- Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously.
- Strong Python and solid engineering practices (reproducibility, CI/CD).
- Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics.
- Familiarity with LLM eval frameworks and LLM-as-judge patterns.
Nice to have
- Experience evaluating agentic / multi-step LLM systems.
- Familiarity with RAG, structured output, and managed LLMs in-VPC.
- FinTech / financial-services domain or other high-stakes, correctness-critical AI.
- Background in statistics or measurement / metrics design.
Worth weighing
- The role involves significant responsibility in defining quality standards for AI outputs, which may lead to high pressure.
- The technical stack includes advanced AI evaluation frameworks, which may require continuous learning and adaptation.
- No salary or specific benefits mentioned in the posting.
Summarised from Phizenix's posting. Read the full original.
Listed by Phizenix on their greenhouse job board, last confirmed open on 2026-09-17. PitchMeAI is not the employer.
More roles at Phizenix
- QA Automation EngineerHyderabad, INDIA (Hybrid)
- Sr. Staff Backend EngineerHyderabad, INDIA (Hybrid)
- Principal AI EngineerMidtown Manhattan, New York City
- Principal Full Stack EngineerMidtown Manhatten, NY (Hybrid)