Martian
A model-routing interface layer for evaluating benchmark performance across many LLMs, plus ownership of the customer-facing sites and portals for an AI infrastructure company.
Making an N-way model comparison legible
Martian routes each request to the model best suited to it on cost, quality and reliability, which only works if the routing decision rests on real comparative measurement. That produces a genuinely awkward interface problem: benchmark results across many models and many providers, on axes that trade against each other rather than resolving to one number.
Presenting that badly reduces it to a leaderboard, which hides the tradeoff that makes routing valuable in the first place. Presenting it well means letting an engineer see why a cheaper model was the correct call for a given class of prompt.
An interface built around the tradeoff, not around a ranking
We built the evaluation surface to keep cost, latency and quality visible together, and owned the public and customer-facing properties around it so the story stayed consistent from marketing site to product.
The surface for evaluating candidate models against one another, structured so comparisons hold across differing provider shapes.
Rendering performance data across many models without collapsing it into a single misleading ranking.
The authenticated product surfaces, owned end to end in React and TypeScript.
The marketing properties alongside the product, kept on the same component foundation to avoid divergence.
Owned the surfaces customers actually touch
Rainier owned and built the customer-facing sites and portals, and the interface layer used to evaluate model performance, for a company whose routing work is backed by published benchmark research.
Have a system with the same constraints?
Send us the scope and the compliance requirements. We will tell you what the architecture should look like before you commit to anything.