Machine Learning Research Scientist, Evaluations
Scale AI · San Francisco, CA; Seattle, WA; New York, NY
Posted about a month ago
or apply directly on Scale AI's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 42 days Scale AI's roles stay open a median of 66 days
- Reposted
- No
- Salary listed
- No 2% of Scale AI's roles list one
- Ghost-job risk at Scale AI
- high 192 stale, 5 reposted of 217 open
- Hiring momentum
- 270 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-10-08
Measured from postings appearing on and disappearing from Scale AI's own greenhouse board since 2026-08-03. Full hiring picture for Scale AI.
About this role
In this role as a Machine Learning Research Scientist focused on evaluations, you will analyze model behavior to identify and diagnose failure modes in large language models (LLMs) and agents. You will design benchmarks and evaluation methods for both text and multimodal modalities, apply post-training techniques, and publish your research findings in top-tier AI conferences. Collaboration with researchers and engineers will be key to defining best practices in evaluation-driven AI development.
- benefits
- 3/5
- freshness
- 4/5
- career value
- 5/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Scale AI as an employer.
What you need
- Expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning
- Experience with LLM evaluation or benchmark development
- Excellent written and verbal communication skills
Nice to have
- Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals
- Previous experience in a customer facing role
What you get
- Comprehensive health, dental and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend (may be eligible)
Worth weighing
- Salary range provided is broad and may vary based on multiple factors
- 90-day waiting period for reconsideration of candidates for the same role
- Role involves significant collaboration and may require strong interpersonal skills
Summarised from Scale AI's posting. Read the full original.
Listed by Scale AI on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.
More roles at Scale AI
- Staff Frontier Agents Engineer (Applied AI)San Francisco, CA; New York, NY
- Technical Program Manager, Public SectorWashington, DC
- Engagement Manager, Public SectorWashington, DC
- Communications Senior Manager, Corporate & Product (Enterprise)San Francisco, CA; New York, NY
- Solutions Engineer, EnterpriseSan Francisco, CA; New York, NY
- Machine Learning Fellow - Human Frontier Collective (US)United States
- Staff Solutions Engineer, EnterpriseNew York, NY; San Francisco, CA
- STEM Fellow - Human Frontier Collective (US)United States