Cortex

Benchmarks

Model Tuning starts with private evals on held-out work. Public benchmarks publish reproducible measurements of Cortex Smart Routing quality and cost without substituting for your real traffic.

Benchmark

Cortex Router on tau2-bench

A first-turn routing counterfactual on tau2-bench.

  • 247/276 tasks passed (89.5%).
  • $0.022 per task for Cortex OSS.
  • Includes train and test tasks.
Benchmark

Cortex Router vs OpenRouter

Claude-tier routing across 400 prompts.

  • 91% of fixed Opus quality.
  • 40% lower normalized cost per prompt vs fixed Opus.
  • Same study as Not Diamond.
Benchmark

Cortex Router vs Sol

Routing quality and cost on 500 SWE-bench Verified tasks.

  • 75.2% vs 79.2% quality credit (Cortex vs Sol).
  • 27% lower model cost per task vs Sol.
Benchmark

Cortex Router vs Not Diamond

Claude-tier routing across 400 prompts.

  • 91% of fixed Opus quality.
  • 40% lower normalized cost per prompt vs fixed Opus.
  • Same study as OpenRouter.