Nextcraft

AI Safety Researcher

Anthropic Research
San Francisco, CA$180K–$280KSeniorPosted 26 days ago
88%

Job description

Research alignment, interpretability, and robustness of frontier models. Design experiments, run evaluations, and publish findings that improve the safety of deployed AI systems.

Anthropic Research is looking for a Senior-level practitioner who can own outcomes end-to-end. You will partner with product, research, and platform teams to ship reliable AI-powered features, instrumented with the evaluation harnesses this role is responsible for.

Day-to-day you will design systems around Alignment, Interpretability, Red teaming, write production code, review pull requests, run evaluations, and contribute to a blameless postmortem culture. We expect strong written communication and comfort working in a fast-moving, evidence-driven environment.

This role is on-site in San Francisco, CA and reports into the AI Safety practice. Compensation ranges from $180K–$280K plus equity and benefits. We are an equal-opportunity employer and actively seek candidates from non-traditional backgrounds who have built real AI systems.

Required competencies

  • Safety · Capability C001
  • Safety · Capability C007
  • Safety · Capability C002

AI-Matched Skills

How your competency profile maps to this role, computed by Nextcraft's match engine.

Alignment88%
Interpretability84%
Red teaming80%
Python76%
Research72%
Overall match score 88%

Employer

Anthropic ResearchAI Safety
Size
1,000+
Location
San Francisco, CA