Mercor
Subcontracted delivery across AI programs, including model training work and evaluation and valuation tooling for an expert-data marketplace.
Evaluating models when the ground truth is human expertise
Mercor's business rests on expert human judgement: instead of low-cost annotation, it puts doctors, lawyers, engineers and bankers in front of model outputs to assess quality in domains where a non-expert cannot tell good from plausible.
That makes the surrounding engineering unusual. The workflows have to route the right expert to the right task, capture judgement in a form models can train against, and hold up as the volume of evaluation work grows quickly.
Support the training and evaluation loop
We worked across several AI programs rather than a single product, contributing to the model training work and to the tooling used to evaluate and value model output.
Delivery against training programs, working to the specifications set by the labs the work served.
Tooling supporting the assessment of model output against expert judgement.
Supporting the valuation side of the evaluation pipeline.
Working across multiple concurrent AI programs rather than a single fixed scope.
Delivery inside a fast-scaling AI data business
Rainier delivered across several AI programs for Mercor, whose expert-evaluation marketplace has scaled to serve the major AI labs.
Have a system with the same constraints?
Send us the scope and the compliance requirements. We will tell you what the architecture should look like before you commit to anything.