or apply directly on Anthropic's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 12 days Anthropic's roles stay open a median of 40 days
- Reposted
- No
- Salary listed
- No 0% of Anthropic's roles list one
- Ghost-job risk at Anthropic
- high 276 stale, 19 reposted of 603 open
- Hiring momentum
- 795 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-09-17
Measured from postings appearing on and disappearing from Anthropic's own greenhouse board since 2026-08-03. Full hiring picture for Anthropic.
About this role
As a Product Designer focused on Evals and Prompts at Anthropic, you will be responsible for writing and revising prompts for AI tools, testing product surfaces, and developing evaluation tools that enable designers to assess and improve prompts. You will collaborate with surface owners and engineers, support model releases, and ensure the alignment of AI behavior with user expectations and safety requirements.
- benefits
- 4/5
- freshness
- 4/5
- career value
- 4/5
- role clarity
- 4/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Anthropic as an employer.
What you need
- Production-quality Python
- Experience building and maintaining evaluation pipelines for LLM products: graders, rubrics, comparison sets, regression suites, and the plumbing that runs them across models
- Experience building internal tools with a real interface for people who do not write code
- Experience standing up test harnesses, sandboxing tool calls, and pinning the settings that make runs comparable
- Experience shipping prompts, or working closely with people who do, and understanding why a prompt that works on one model fails on the next
- Reads transcripts, not only scores
Nice to have
- Has worked inside a model-launch cycle
- A/B testing experience and the ability to connect offline evals to online outcomes
- Front-end or notebook-to-app experience, and opinions about what makes an eval result legible at a glance
- Has turned product rubrics into training signal: graders, human-feedback questions, or preference pairs
- Cares how Claude behaves for the people using it, not only whether the metric moved
What you get
- Annual Salary: $305,000 — $385,000 USD
- Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time
- Visa sponsorship available
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
Worth weighing
- No specific mention of the tech stack used in the role
- Role involves significant collaboration with multiple teams, which may require strong communication skills
- The position may involve working with legacy systems or processes, depending on the existing evaluation pipelines
Summarised from Anthropic's posting. Read the full original.
Listed by Anthropic on their greenhouse job board, last confirmed open on 2026-09-17. PitchMeAI is not the employer.
More roles at Anthropic
- Enterprise Integrated Campaign ManagerSan Francisco, CA | New York City, NY | Seattle, WA
- Finance Systems Engineer, Finance and StrategySan Francisco, CA
- Strategic Account Executive, IndustriesSan Francisco, CA | New York City, NY
- Strategy & Operations Lead, Enterprise MarketingSan Francisco, CA | New York City, NY
- Customer Marketing Manager, IndustriesSan Francisco, CA | New York City, NY
- Staff+ Software Engineer, Platform EcosystemSan Francisco, CA | New York City, NY
- Software Engineering Manager, Network SecuritySan Francisco, CA | New York City, NY
- Head of Treasury Strategy & TransformationRemote-Friendly (Travel Required) | San Francisco, CA