Job description
Optimize and deploy models to edge hardware. Quantize, distill, and compile models for low-latency inference on field devices with constrained budgets.
Replicate is looking for a Mid-level practitioner who can own outcomes end-to-end. You will partner with product, research, and platform teams to ship reliable AI-powered features, instrumented with the evaluation harnesses this role is responsible for.
Day-to-day you will design systems around Edge ML, TensorRT, C++, write production code, review pull requests, run evaluations, and contribute to a blameless postmortem culture. We expect strong written communication and comfort working in a fast-moving, evidence-driven environment.
This role is remote-friendly and reports into the ML Deployment practice. Compensation ranges from $130K–$195K plus equity and benefits. We are an equal-opportunity employer and actively seek candidates from non-traditional backgrounds who have built real AI systems.
Required competencies
- Operator · Capability C008
- Orchestration · Capability C012
- Orchestration · Capability C014
AI-Matched Skills
How your competency profile maps to this role, computed by Nextcraft's match engine.