Job description
Research alignment, interpretability, and robustness of frontier models. Design experiments, run evaluations, and publish findings that improve the safety of deployed AI systems.
Anthropic Research is looking for a Senior-level practitioner who can own outcomes end-to-end. You will partner with product, research, and platform teams to ship reliable AI-powered features, instrumented with the evaluation harnesses this role is responsible for.
Day-to-day you will design systems around Alignment, Interpretability, Red teaming, write production code, review pull requests, run evaluations, and contribute to a blameless postmortem culture. We expect strong written communication and comfort working in a fast-moving, evidence-driven environment.
This role is on-site in San Francisco, CA and reports into the AI Safety practice. Compensation ranges from $180K–$280K plus equity and benefits. We are an equal-opportunity employer and actively seek candidates from non-traditional backgrounds who have built real AI systems.
Required competencies
- Safety · Capability C001
- Safety · Capability C007
- Safety · Capability C002
AI-Matched Skills
How your competency profile maps to this role, computed by Nextcraft's match engine.