Nextcraft

Evaluation Engineer

Cohere
Toronto, CanadaRemote$120K–$180KMidPosted 8 days ago
79%

Job description

Build and operate the eval platform for LLM-powered products. Design LLM-as-judge pipelines, regression suites, and human-eval panels for production models.

Cohere is looking for a Mid-level practitioner who can own outcomes end-to-end. You will partner with product, research, and platform teams to ship reliable AI-powered features, instrumented with the evaluation harnesses this role is responsible for.

Day-to-day you will design systems around LLM evaluation, Python, Statistics, write production code, review pull requests, run evaluations, and contribute to a blameless postmortem culture. We expect strong written communication and comfort working in a fast-moving, evidence-driven environment.

This role is remote-friendly and reports into the Enterprise LLMs practice. Compensation ranges from $120K–$180K plus equity and benefits. We are an equal-opportunity employer and actively seek candidates from non-traditional backgrounds who have built real AI systems.

Required competencies

  • Orchestration · Capability C007
  • Safety · Capability C002

AI-Matched Skills

How your competency profile maps to this role, computed by Nextcraft's match engine.

LLM evaluation79%
Python75%
Statistics71%
Data labeling67%
Overall match score 79%

Employer

CohereEnterprise LLMs
Size
400+
Location
Toronto, Canada