Every prompt is a buying decision
Choosing one model for every request makes deployment simple, but it also commits the same level of capability to very different work. A short extraction task and a difficult reasoning problem may justify different choices.
Cortex makes that choice before the answer is generated. It uses the prompt to select a model tier; this study examines the reported answer quality and route-weighted cost of those decisions.
What we compared
Cortex and OpenRouter selected among the same three Claude tiers on 400 test prompts. Fixed-model policies send every prompt to Opus, Sonnet, or Haiku. The reported quality score is the mean judged quality of the selected answers on a 0–1 scale; it is distinct from matching a reference route.
Cortex was trained on this traffic, while the off-the-shelf comparator had different prior exposure. The OpenRouter and Not Diamond pages present the same archived study records, not independent experiments. These differences limit what the comparison can establish beyond this cohort.
Overview
- 0.803
- Reported quality · 0–1+0.125 vs OpenRouter
- 0.678
- Reported quality · 0–1−0.125 vs Cortex
- 0.882
- 91.0% retained by Cortex
- 39.5%
- Normalized units vs Opus
Quality and cost
Cortex’s reported quality of 0.803 falls below fixed Opus (0.882) and above OpenRouter (0.678). The dollar illustration below compares historical input prices; the sweep compares separate study cost units. Neither establishes equal measured bills.
Reported quality · higher is better
Historical blended input $/MTok · lower is cheaper
| Fixed | 0.504 | $1.000 | One model for every prompt |
| Fixed | 0.575 | $3.000 | One model for every prompt |
| Oracle (hindsight) | 0.927 | $3.570 | 400 reference-tier labels |
| 0.803 | $3.845 | 400 known routes | |
| 0.678 | $4.008 | 393 known routes; unresolved routes excluded | |
| Fixed | 0.882 | $5.000 | One model for every prompt |
- Historical input prices assume equal token volume per known route; they do not include measured token lengths or routing overhead.
- OpenRouter's price is averaged over 393 known routes: $4.008 / MTok. Unresolved routes are excluded.
- The Oracle price applies historical rates to the reference-tier mix; it is a hindsight illustration, not an available routing policy.
| Fixed | 0.0400 | 1 unit per prompt |
| Fixed | 0.2000 | 5 units per prompt |
| Oracle (hindsight) | 0.529 | Reported reference point; no deployable oracle |
| 0.5925 | 393 routes; 7 unresolved | |
| 0.605239.5% lower vs fixed Opus | 400 recorded routes | |
| Fixed | 1.0000 | 25 units per prompt |
Reported cost–quality settings
Reported quality · higher is better
Normalized cost per prompt · fixed Opus = 1
| 0 | 0.867 | 0.980 | 98.0% |
| 1 | 0.678 | 0.593 | 49.5% |
| 2 | 0.563 | 0.201 | 0.8% |
| 3 | 0.497 | 0.042 | 0.2% |
| 4 | 0.496 | 0.042 | 0.2% |
| 5 | 0.495 | 0.039 | 0.0% |
| 6 | 0.495 | 0.039 | 0.0% |
| 7 | 0.495 | 0.039 | 0.0% |
| 8 | 0.495 | 0.039 | 0.0% |
| 9 | 0.495 | 0.039 | 0.0% |
| 10 | 0.495 | 0.039 | 0.0% |
| 0.040 | 0.504 |
| 0.250 | 0.675 |
| 0.500 | 0.787 |
| 0.605 | 0.803 |
| 0.750 | 0.845 |
| 1.000 | 0.882 |
Budget-curve area and operating guidance
| 0.790 | |
| Random routing | 0.729 |
| Length heuristic | 0.717 |
| 0.681 |
- The archive calls APGR the area under the quality-versus-budget curve. The complete area calculation and length/random curves are unavailable; these are reported comparisons, not newly verified rankings.
- The comparator has only a few distinct operating points. Joining them does not demonstrate a reachable setting at every budget.
- The original report also described 71% of the Oracle gap closed for Cortex versus 41% for OpenRouter. Its stated reference baseline is inconsistent; those percentages are archived claims, not a recomputed metric.
| 95% of always- | 0.838 | 0.707 |
| 98% of always- | 0.864 | 0.876 |
- Choose an operating point by setting a quality floor, a cost budget, or the value of an additional quality improvement. The tables preserve the report's examples; they are not production guarantees.
- Reported budget examples: 0.500 study cost yields 0.787 quality; 0.750 yields 0.845. The stored quality targets report 95% of Opus quality at 0.707 cost and 98% at 0.876.
- The original discussion suggested a knee near 35–55% Opus use. That range is guidance from the archive, not an optimum established by these sparse samples; Opus call share and normalized cost are different measures.
- Re-evaluate a selected threshold on the intended workload before drawing conclusions about reachable quality or savings.
Where each prompt was routed
- Cortex routes 211 prompts to Opus and 42 to Haiku.
- Router agreement: 144 / 400 prompts (36.0%).
- Missing comparator routes count as disagreements.
- Optimal-route rate is reported as 0.760 for Cortex and 0.620 for OpenRouter: selecting a best-scoring model is distinct from exact-label accuracy. The raw scores needed to recompute this rate are unavailable.
- Precision, recall and F1 use the recorded confusion counts. OpenRouter's precision / recall therefore display 0.323 / 0.378, rather than the archived rounded values 0.324 / 0.377.
| 245 / 400 | 61.25% | 0.609 | 0.613 | 0.605 | |
| 151 / 400 | 37.75% | 0.323 | 0.378 | 0.348 |
- Exact-label accuracy: selected tier matches ground truth; distinct from judged answer quality.
- Precision, recall, and F1: per-tier metrics weighted by support.
- OpenRouter, excluding missing routes: 151 / 393 (38.42%).
Routing accuracy by model
- Rows: ground-truth tiers; columns: selected tiers.
- OpenRouter original matrix: 393 recorded routes.
- Missing column = ground-truth support minus recorded routes in each row.
| 0.597 | 0.712 | 0.649 | 0.374 | 0.418 | 0.395 | 177 | |
| 0.673 | 0.619 | 0.645 | 0.395 | 0.481 | 0.434 | 160 | |
| 0.476 | 0.317 | 0.381 | 0.000 | 0.000 | 0.000 | 63 |
- Precision is displayed as zero for tiers with no predictions.
- P = precision; R = recall. Tiers are sorted by support, highest first.
- Large-tier recall is an under-routing risk probe: Cortex 0.712 versus OpenRouter 0.418. It measures how often prompts labeled Opus were routed to Opus.
- Small-tier precision asks whether a cheap route matches its reference tier: Cortex 0.476. OpenRouter makes no Haiku predictions; its displayed zero is not evidence about the quality of Haiku answers.
- Smallest ground-truth tier: 63 prompts; less support than the larger tiers.
Where the routers agree
- Missing column inferred from Cortex route totals and recorded route pairs.
- The largest split decisions exchange Opus and Sonnet: 126 prompts go from Cortex Opus to comparator Sonnet, and 83 in the other direction.
- Of Cortex's 42 Haiku routes, 30 receive comparator Opus, 10 comparator Sonnet, and 2 have no comparator route.
Inspect individual routing calls
- 16 selected routing excerpts; not the complete test set.
- Original prompt lengths preserved; full prompts and model responses unavailable.
- In this selected sample, Cortex matches 13 reference tiers and OpenRouter matches 4. These sample counts do not estimate full-cohort accuracy.
Showing 16 of 16 excerpts.
| No route · Unresolved |
Prompt #0
Excerpt from a 15,919-character prompt.
...levant to your task. </system-reminder> start a worktree to solve 19 private _ functions imported from provider_endpoint_inventory │ Important │ export_scrubbed_fixture.py:29–50 │ Reviewer says follow-up PR acceptable
| Ground truth | Reference label | |
| Yes | ||
| No |
- Reported classifier probabilities: L 0.85 · M 0.14 · S 0.01.
- L / M / S identify Opus / Sonnet / Haiku.
Methodology and limits
The archived report describes prompts from work such as coding tasks, PR reviews, agent traces, and safety-rule writing. In its reported pipeline, each prompt becomes an embedding; a lightweight classifier scores the model tiers and applies the routing rule before answer generation.
| Dataset | 400 evaluation prompts from work traffic |
| Split | 600 training / 400 test prompts; reported overlap 0 |
| Embedding model | |
| Classifier | Random forest; no class weighting |
| Routing rule | Highest predicted model-tier score; reported threshold 0.5 |
| Evaluation | Frozen judge score matrix shared by the compared policies |
- Cortex trained on this traffic; the off-the-shelf comparator had different prior exposure, limiting generalization.
- Original split, score matrix, and complete prompt-level routes unavailable for independent verification.
- Session caching and production token accounting were not measured.
Archived work-memory embedding projection

Routing with the cache in mind
The original report used this illustration to explain why a single-turn routing saving may disappear in a longer agent session. A reusable prefix can include system instructions, tool definitions and conversation context; its cached input price differs from the cost of processing it again.
- A cache-blind router can save on a model's list rate while paying more to reprocess a long prefix. Compare the avoided cache benefit with the expected quality gain and cost of a switch.
- The archived cache-aware routing discussion calls for tracking warm state, expiry and prefix changes; preserving a useful main-session cache; and considering cheaper subagents for narrower tasks. These production behaviors were not evaluated in this benchmark.
| Short conversations | Little reusable context accumulates; choose primarily for the required quality. |
| Strong model plus cheaper models | Compare keeping the strong model's warm prefix with the quality and price benefit of moving a turn. |
| Cheap-only model pool | When token-price differences are small, quality may matter more than preserving the cache. |
The archive's example economics assume cache rates; actual discounts, retention and routing overhead require workload-specific measurement. The study contains no session-level token accounting.
Archived list-price illustration
| $1.000 | $5.000 | 42 | 0 | ||
| $3.000 | $15.000 | 147 | 195 | ||
| $5.000 | $25.000 | 211 | 198 |
| 400 | $3.845 | $19.225 | |
| 393 | $4.008 | $20.038 |
- Model list rates are sorted from lowest to highest input price; the route columns retain each tier's recorded counts.
- Historical rates assume equal token volume per route; unknown routes excluded.
- Excludes token-length variation, cache effects, batching, and router overhead.
- List prices use a different cost basis from study units; this illustration is not a measured bill.