Benchmarks
Model Tuning starts with private evals on held-out work. Public benchmarks publish reproducible measurements of Cortex Smart Routing quality and cost without substituting for your real traffic.
Benchmark
Cortex Router on tau2-bench
A first-turn routing counterfactual on tau2-bench.
- 247/276 tasks passed (89.5%).
- $0.022 per task for Cortex OSS.
- Includes train and test tasks.
Benchmark
Cortex Router vs OpenRouter
Claude-tier routing across 400 prompts.
- 91% of fixed Opus quality.
- 40% lower normalized cost per prompt vs fixed Opus.
- Same study as Not Diamond.
Benchmark
Cortex Router vs Sol
Routing quality and cost on 500 SWE-bench Verified tasks.
- 75.2% vs 79.2% quality credit (Cortex vs Sol).
- 27% lower model cost per task vs Sol.
Benchmark
Cortex Router vs Not Diamond
Claude-tier routing across 400 prompts.
- 91% of fixed Opus quality.
- 40% lower normalized cost per prompt vs fixed Opus.
- Same study as OpenRouter.