Job description
Build and operate the eval platform for LLM-powered products. Design LLM-as-judge pipelines, regression suites, and human-eval panels for production models.
Cohere is looking for a Mid-level practitioner who can own outcomes end-to-end. You will partner with product, research, and platform teams to ship reliable AI-powered features, instrumented with the evaluation harnesses this role is responsible for.
Day-to-day you will design systems around LLM evaluation, Python, Statistics, write production code, review pull requests, run evaluations, and contribute to a blameless postmortem culture. We expect strong written communication and comfort working in a fast-moving, evidence-driven environment.
This role is remote-friendly and reports into the Enterprise LLMs practice. Compensation ranges from $120K–$180K plus equity and benefits. We are an equal-opportunity employer and actively seek candidates from non-traditional backgrounds who have built real AI systems.
Required competencies
- Orchestration · Capability C007
- Safety · Capability C002
AI-Matched Skills
How your competency profile maps to this role, computed by Nextcraft's match engine.