Research Engineer, Model Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY
Posted about 2 months ago
or apply directly on Anthropic's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 66 days Anthropic's roles stay open a median of 62 days
- Reposted
- No
- Salary listed
- No 0% of Anthropic's roles list one
- Ghost-job risk at Anthropic
- high 449 stale, 19 reposted of 603 open
- Hiring momentum
- 795 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-10-08
Measured from postings appearing on and disappearing from Anthropic's own greenhouse board since 2026-08-03. Full hiring picture for Anthropic.
About this role
In this role, you will lead the evaluation strategy for production Claude models, focusing on optimizing measurement techniques to assess model quality. You will monitor the evaluation fleet during live training runs, provide guidance on evaluation methodologies, and investigate any regressions in production runs. Additionally, you will build dashboards and reports for tracking model performance, ensuring that the evaluation process is both rigorous and trustworthy.
- benefits
- 3/5
- freshness
- 1/5
- career value
- 4/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Anthropic as an employer.
What you need
- Designed, run, and analyzed evaluations for ML models at scale
- Strong Python skills and comfort with production systems
- Ability to turn ambiguous results into clear recommendations
- Maintain clarity and rigor when debugging complex, time-sensitive issues
- Ability to thrive in controlled chaos and manage multiple urgent priorities
- Care about the societal impacts of work and responsible shipping of frontier models
Nice to have
- Hands-on experience post-training large language models
- Research experience in ML evaluation or benchmarking
- Background in statistics and experimental design
- Experience developing robust evaluation metrics for ML systems
What you get
- Annual Salary: $500,000 — $850,000 USD
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Lovely office space for collaboration
Worth weighing
- No specific mention of required years of experience or exact educational background
- Visa sponsorship is available but not guaranteed for every role or candidate
- The role may require being in the office more than 25% of the time depending on specific job requirements
Summarised from Anthropic's posting. Read the full original.
Listed by Anthropic on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.
More roles at Anthropic
- Staff Software Engineer, Environments InfrastructureSan Francisco, CA | New York City, NY
- Research Engineer, Cybersecurity RL (Reinforcement Learning)San Francisco, CA | New York City, NY
- Technical Program Manager, API PlatformSan Francisco, CA | Seattle, WA
- Senior Manager, Technical Accounting - M&A and InvestmentsSan Francisco, CA | Seattle, WA
- Product Operations Manager, EmbeddedSan Francisco, CA | New York City, NY | Seattle, WA
- Staff Engineer, Datacenter Server LifecycleSydney, Australia
- Staff+ Software Engineer, Capacity EngineeringSan Francisco, CA | New York City, NY | Seattle, WA
- Applied AI Architect, PartnershipsSan Francisco, CA | New York City, NY