Machine Learning Research Scientist, Evaluations
Scale AI · San Francisco, CA; Seattle, WA; New York, NY
Posted 21 days ago
or apply directly on Scale AI's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 21 days Scale AI's roles stay open a median of 45 days
- Reposted
- No
- Salary listed
- No 2% of Scale AI's roles list one
- Ghost-job risk at Scale AI
- low 0 stale, 5 reposted of 217 open
- Hiring momentum
- 270 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-09-17
Measured from postings appearing on and disappearing from Scale AI's own greenhouse board since 2026-08-03. Full hiring picture for Scale AI.
About this role
In this role as a Machine Learning Research Scientist focused on evaluations, you will analyze model behavior to identify and diagnose failure modes in large language models (LLMs) and agents. You will design benchmarks and evaluation methods for both text and multimodal modalities, apply post-training techniques, and publish your research findings in top-tier AI conferences. Collaboration with researchers and engineers will be key to defining best practices in evaluation-driven AI development.
- benefits
- 3/5
- freshness
- 4/5
- career value
- 5/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Scale AI as an employer.
What you need
- Expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning
- Experience with LLM evaluation or benchmark development
- Excellent written and verbal communication skills
Nice to have
- Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals
- Previous experience in a customer facing role
What you get
- Comprehensive health, dental and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend (may be eligible)
Worth weighing
- Salary range provided is broad and may vary based on multiple factors
- 90-day waiting period for reconsideration of candidates for the same role
- Role involves significant collaboration and may require strong interpersonal skills
Summarised from Scale AI's posting. Read the full original.
Listed by Scale AI on their greenhouse job board, last confirmed open on 2026-09-17. PitchMeAI is not the employer.
More roles at Scale AI
- Staff Product DesignerNew York, NY; Washington, DC
- Strategic Finance Manager, CorporateSan Francisco, CA
- Partner Development ManagerWashington, DC
- Enterprise Deal Desk AnalystSan Francisco, CA; New York, NY
- Senior Software Engineer, AI Operations, GPSDoha, Qatar
- Engagement Manager, Public Sector (Midwest)Omaha, NE
- Data Acquisition Lead, Frontier EnvironmentsSan Francisco, CA; New York, NY
- Program Manager, ComplianceSan Francisco, CA