Anthropic

Safeguards Enforcement Analyst, Violence & Extremism

Anthropic · Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Posted about 2 months ago

or apply directly on Anthropic's site. We never take the application ourselves.

Is this posting real?

This role has been open
66 days
Anthropic's roles stay open a median of 62 days
Reposted
No
Salary listed
No
0% of Anthropic's roles list one
Ghost-job risk at Anthropic
high
449 stale, 19 reposted of 603 open
Hiring momentum
795 roles opened in the last 90 days
↑ up vs. the prior 90 days
Last confirmed on the employer's board
2026-10-08

Measured from postings appearing on and disappearing from Anthropic's own greenhouse board since 2026-08-03. Full hiring picture for Anthropic.

About this role

As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will design and implement operational workflows to assess AI model behavior and enforce policies against misuse that could lead to real-world harm. Your role involves collaborating with engineering and data science teams to optimize detection systems, reviewing flagged content, and staying informed about emerging threats and extremist movements. You will also develop guidelines and provide feedback on policy gaps based on enforcement experiences.

Our read on this posting2.4out of 5
benefits
3/5
freshness
1/5
career value
4/5
role clarity
4/5
pay transparency
0/5

Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Anthropic as an employer.

What you need

  • Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
  • Experience standing up and scaling policy enforcement or content review workflows
  • Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
  • Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement
  • Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space

Nice to have

  • Subject matter expertise in one or more high-stakes harm areas, such as weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
  • Familiarity with relevant legal and regulatory frameworks governing dangerous technology, critical infrastructure, or domestic/international terrorism
  • Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases
  • Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
  • Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments

What you get

  • Annual Salary: $285,000 — $330,000 USD
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Lovely office space for collaboration

Worth weighing

  • Exposure to explicit content that may be violent, graphic, or psychologically disturbing
  • Visa sponsorship is available, but not guaranteed for every candidate
  • The role may require in-office presence at least 25% of the time, depending on specific role requirements

Summarised from Anthropic's posting. Read the full original.

Listed by Anthropic on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.

More roles at Anthropic

All 603 open roles at Anthropic →