← Past performance
Case study 10AI · Model evaluation · Subcontractor

Mercor

Subcontracted delivery across AI programs, including model training work and evaluation and valuation tooling for an expert-data marketplace.

The challenge

Evaluating models when the ground truth is human expertise

Mercor's business rests on expert human judgement: instead of low-cost annotation, it puts doctors, lawyers, engineers and bankers in front of model outputs to assess quality in domains where a non-expert cannot tell good from plausible.

That makes the surrounding engineering unusual. The workflows have to route the right expert to the right task, capture judgement in a form models can train against, and hold up as the volume of evaluation work grows quickly.

Our approach

Support the training and evaluation loop

We worked across several AI programs rather than a single product, contributing to the model training work and to the tooling used to evaluate and value model output.

01
Model training work

Delivery against training programs, working to the specifications set by the labs the work served.

02
Evaluation tooling

Tooling supporting the assessment of model output against expert judgement.

03
Valuation work

Supporting the valuation side of the evaluation pipeline.

04
Delivery across programs

Working across multiple concurrent AI programs rather than a single fixed scope.

Outcome

Delivery inside a fast-scaling AI data business

Rainier delivered across several AI programs for Mercor, whose expert-evaluation marketplace has scaled to serve the major AI labs.

Role
Subcontract
multiple programs
Domain
AI evaluation
training and assessment
Scope
Multi-program
concurrent workstreams

Have a system with the same constraints?

Send us the scope and the compliance requirements. We will tell you what the architecture should look like before you commit to anything.