Cortex
Skip to benchmark report

Cortex Router vs OpenRouter

Claude Routing: Quality and Cost

Every prompt is a buying decision

Choosing one model for every request makes deployment simple, but it also commits the same level of capability to very different work. A short extraction task and a difficult reasoning problem may justify different choices.

Cortex makes that choice before the answer is generated. It uses the prompt to select a model tier; this study examines the reported answer quality and route-weighted cost of those decisions.

What we compared

Cortex and OpenRouter selected among the same three Claude tiers on 400 test prompts. Fixed-model policies send every prompt to Opus, Sonnet, or Haiku. The reported quality score is the mean judged quality of the selected answers on a 0–1 scale; it is distinct from matching a reference route.

Cortex was trained on this traffic, while the off-the-shelf comparator had different prior exposure. The OpenRouter and Not Diamond pages present the same archived study records, not independent experiments. These differences limit what the comparison can establish beyond this cohort.

Overview

0.803
Reported quality · 0–1+0.125 vs OpenRouter
0.678
Reported quality · 0–1−0.125 vs Cortex
0.882
91.0% retained by Cortex
39.5%
Normalized units vs Opus

Quality and cost

Cortex’s reported quality of 0.803 falls below fixed Opus (0.882) and above OpenRouter (0.678). The dollar illustration below compares historical input prices; the sweep compares separate study cost units. Neither establishes equal measured bills.

Reported selected-model quality (0–1)
01
Legend
  • Oracle (hindsight)
    Best model per prompt, chosen after answer scores are known.
  • Fixed Opus 4.6
    Every prompt uses Opus 4.6.
  • Cortex
    Selects a model tier for each prompt.
  • OpenRouter
    Selects a model tier for each prompt.
  • Fixed Sonnet 4.6
    Every prompt uses Sonnet 4.6.
  • Fixed Haiku 4.5
    Every prompt uses Haiku 4.5.
  • Oracle ceiling (hindsight)
    Dashed reference at 0.927 on the shared scale.

400 prompts. Quality is the reported mean judged answer score on a 0–1 scale; raw judge scores are unavailable.

OpenRouter has 7 unresolved routes. Their treatment in the quality average is unverified.

Reported quality (0–1). Bars are ordered highest first.

Hover, focus or select a bar for details.

Quality versus historical input price

Reported quality · higher is better

Historical blended input $/MTok · lower is cheaper

Legend
  • Haiku 4.5
    Every prompt uses Haiku 4.5.
  • Sonnet 4.6
    Every prompt uses Sonnet 4.6.
  • Oracle · Claude tiers
    Hindsight ceiling across Claude tiers; not a deployable router.
  • Cortex
    Reported operating point; the router chooses among the three Claude tiers.
  • OpenRouter
    Reported operating point; the router chooses among the three Claude tiers.
  • Opus 4.6
    Every prompt uses Opus 4.6.

Historical list rates assume equal input tokens per known route; this is not a measured bill and excludes routing overhead.

OpenRouter: 393 known routes; 7 unresolved routes excluded. Oracle input price uses the ground-truth tier mix. Select a logo for exact values.

Historical input-price comparison · lowest price first
Historical input-price comparison · lowest price first
Fixed Haiku 4.50.504$1.000One model for every prompt
Fixed Sonnet 4.60.575$3.000One model for every prompt
Oracle (hindsight)0.927$3.570400 reference-tier labels
Cortex0.803$3.845400 known routes
OpenRouter0.678$4.008393 known routes; unresolved routes excluded
Fixed Opus 4.60.882$5.000One model for every prompt
  • Historical input prices assume equal token volume per known route; they do not include measured token lengths or routing overhead.
  • OpenRouter's price is averaged over 393 known routes: $4.008 / MTok. Unresolved routes are excluded.
  • The Oracle price applies historical rates to the reference-tier mix; it is a hindsight illustration, not an available routing policy.
Normalized cost per prompt: fixed Opus = 1; lowest cost first
Normalized cost per prompt: fixed Opus = 1; lowest cost first
Fixed Haiku0.04001 unit per prompt
Fixed Sonnet0.20005 units per prompt
Oracle (hindsight)0.529Reported reference point; no deployable oracle
OpenRouter0.5925393 routes; 7 unresolved
Cortex0.605239.5% lower vs fixed Opus400 recorded routes
Fixed Opus1.000025 units per prompt

Cost methodology

Cost basis
Opus / Sonnet / Haiku: 25 / 5 / 1 study units per route.
Calculation
Normalized cost = summed route-weighted units ÷ (400 × 25). Fixed Opus = 1.
Limitations

Study units are not dollars or measured token charges.

OpenRouter: 7 unresolved routes add zero units; all 400 prompts remain in the cost denominator. Unresolved-route charges are unavailable, so comparator operating cost is incomplete.

Reported cost–quality settings

Cost–quality sweep with fixed-policy references

Reported quality · higher is better

Normalized cost per prompt · fixed Opus = 1

Legend
  • Haiku 4.5
    Every prompt uses Haiku 4.5. The horizontal guide marks its reported quality.
  • Sonnet 4.6
    Every prompt uses Sonnet 4.6. The horizontal guide marks its reported quality.
  • Oracle · Claude tiers
    Hindsight ceiling across Claude tiers; not a deployable router. Horizontal guide marks reported quality.
  • OpenRouter
    Dashed guide through archived settings. The larger logo marks the reported operating point.
  • Cortex
    Solid guide through archived settings. The larger logo marks the reported operating point.
  • Opus 4.6
    Every prompt uses Opus 4.6. The horizontal guide marks its reported quality.
  • Cortex · ≈+0.119 quality
    Same cost · estimated: 0.605 × fixed Opus. Cortex 0.803 vs OpenRouter ≈0.684 on the 0–1 quality scale. OpenRouter is linearly interpolated between archived settings. Select the dotted connector for details.

Lines are guides through archived samples; intermediate settings were not measured. Interpolation does not establish reachable settings or measured equal-cost outcomes.

At the highlighted OpenRouter operating point, 7 unresolved routes add zero published cost units; actual charges are unknown. Other settings lack route-level records; raw judge scores are unavailable.

Fixed Opus = 1 in study units, not dollars. Select a logo for exact values and overlapping settings.

OpenRouter: reported tradeoff settings
OpenRouter: reported tradeoff settings
00.8670.98098.0%
10.6780.59349.5%
20.5630.2010.8%
30.4970.0420.2%
40.4960.0420.2%
50.4950.0390.0%
60.4950.0390.0%
70.4950.0390.0%
80.4950.0390.0%
90.4950.0390.0%
100.4950.0390.0%
Cortex: reported curve points
Cortex: reported curve points
0.0400.504
0.2500.675
0.5000.787
0.6050.803
0.7500.845
1.0000.882

Budget-curve area and operating guidance

Budget-curve area (APGR) · reported comparison, highest first
Budget-curve area (APGR) · reported comparison, highest first
Cortex0.790
Random routing0.729
Length heuristic0.717
OpenRouter0.681
  • The archive calls APGR the area under the quality-versus-budget curve. The complete area calculation and length/random curves are unavailable; these are reported comparisons, not newly verified rankings.
  • The comparator has only a few distinct operating points. Joining them does not demonstrate a reachable setting at every budget.
  • The original report also described 71% of the Oracle gap closed for Cortex versus 41% for OpenRouter. Its stated reference baseline is inconsistent; those percentages are archived claims, not a recomputed metric.
Archived Cortex quality targets · lowest required cost first
Archived Cortex quality targets · lowest required cost first
95% of always-Opus0.8380.707
98% of always-Opus0.8640.876
  • Choose an operating point by setting a quality floor, a cost budget, or the value of an additional quality improvement. The tables preserve the report's examples; they are not production guarantees.
  • Reported budget examples: 0.500 study cost yields 0.787 quality; 0.750 yields 0.845. The stored quality targets report 95% of Opus quality at 0.707 cost and 98% at 0.876.
  • The original discussion suggested a knee near 35–55% Opus use. That range is guidance from the archive, not an optimum established by these sparse samples; Opus call share and normalized cost are different measures.
  • Re-evaluate a selected threshold on the intended workload before drawing conclusions about reachable quality or savings.

Where each prompt was routed

Route counts across all 400 prompts
Opus 4.6
Sonnet 4.6
Haiku 4.5
No route
Legend
  • Cortex
    Number of prompts assigned to each tier.
  • OpenRouter
    Number of prompts assigned to each tier.
  • Opus 4.6
    Prompts sent to Opus 4.6; routing share is distinct from answer quality.
  • Sonnet 4.6
    Prompts sent to Sonnet 4.6; routing share is distinct from answer quality.
  • Haiku 4.5
    Prompts sent to Haiku 4.5; routing share is distinct from answer quality.
  • No route
    No usable tier recorded. These prompts remain in the full-cohort denominator.

Higher value first within each group. Select a bar for its definition and value.

Counts include all 400 prompts; unresolved routes shown separately.

  • Cortex routes 211 prompts to Opus and 42 to Haiku.
  • Router agreement: 144 / 400 prompts (36.0%).
  • Missing comparator routes count as disagreements.
Routing diagnostics: four comparisons
Optimal-route rate
Recall
Precision
F1
Legend
  • Cortex
    Router performance for each diagnostic metric.
  • OpenRouter
    Router performance for each diagnostic metric.
  • Optimal-route rate
    Archived share reported as selecting the best-scoring model. Distinct from exact-label accuracy; underlying per-prompt scores are unavailable for recomputation.
  • Recall
    Per-tier recall weighted by support: of prompts with each reference tier, how many were routed there. Missing routes count as misses.
  • Precision
    Per-tier precision weighted by ground-truth support: of routes assigned to each tier, how many matched that tier.
  • F1
    Support-weighted harmonic mean of per-tier precision and recall.

Higher value first within each group. Select a bar for its definition and value.

Precision, recall and F1 recomputed; optimal-route rates are archived aggregates.

  • Optimal-route rate is reported as 0.760 for Cortex and 0.620 for OpenRouter: selecting a best-scoring model is distinct from exact-label accuracy. The raw scores needed to recompute this rate are unavailable.
  • Precision, recall and F1 use the recorded confusion counts. OpenRouter's precision / recall therefore display 0.323 / 0.378, rather than the archived rounded values 0.324 / 0.377.
Exact-label diagnostics, with missing routes counted as misses
Exact-label diagnostics, with missing routes counted as misses
Cortex245 / 40061.25%0.6090.6130.605
OpenRouter151 / 40037.75%0.3230.3780.348
  • Exact-label accuracy: selected tier matches ground truth; distinct from judged answer quality.
  • Precision, recall, and F1: per-tier metrics weighted by support.
  • OpenRouter, excluding missing routes: 151 / 393 (38.42%).

Routing accuracy by model

  • Rows: ground-truth tiers; columns: selected tiers.
  • OpenRouter original matrix: 393 recorded routes.
  • Missing column = ground-truth support minus recorded routes in each row.
Cortex confusion matrix
Cortex confusion matrix
Opus 4.6
Sonnet 4.6
Haiku 4.5
Legend
  • Correct route
    Diagonal cells: the selected tier matches the reference label.
  • Mis-route
    Off-diagonal cells: the selected tier differs from the reference label.
  • No route
    No usable selected tier was recorded. Counts remain in the row denominator.
  • Zero prompts
    Neutral cells display 0 when no prompt has this row and column combination.

Rows: Ground-truth tier. Columns: selected tier. Tiers ordered by study cost; missing routes last.

Cells show prompt counts; deeper color means more prompts. Select a cell for row shares and interpretation.

OpenRouter confusion matrix, including unknown responses
OpenRouter confusion matrix, including unknown responses
Opus 4.6
Sonnet 4.6
Haiku 4.5
Legend
  • Correct route
    Diagonal cells: the selected tier matches the reference label.
  • Mis-route
    Off-diagonal cells: the selected tier differs from the reference label.
  • No route
    No usable selected tier was recorded. Counts remain in the row denominator.
  • Zero prompts
    Neutral cells display 0 when no prompt has this row and column combination.

Rows: Ground-truth tier. Columns: selected tier. Tiers ordered by study cost; missing routes last.

Cells show prompt counts; deeper color means more prompts. Select a cell for row shares and interpretation.

Per-tier precision, recall, and F1 · routers side by side
Per-tier precision, recall, and F1 · routers side by side
Opus 4.60.5970.7120.6490.3740.4180.395177
Sonnet 4.60.6730.6190.6450.3950.4810.434160
Haiku 4.50.4760.3170.3810.0000.0000.00063
  • Precision is displayed as zero for tiers with no predictions.
  • P = precision; R = recall. Tiers are sorted by support, highest first.
  • Large-tier recall is an under-routing risk probe: Cortex 0.712 versus OpenRouter 0.418. It measures how often prompts labeled Opus were routed to Opus.
  • Small-tier precision asks whether a cheap route matches its reference tier: Cortex 0.476. OpenRouter makes no Haiku predictions; its displayed zero is not evidence about the quality of Haiku answers.
  • Smallest ground-truth tier: 63 prompts; less support than the larger tiers.

Where the routers agree

Cortex routes (rows) versus OpenRouter routes (columns)
Cortex routes (rows) versus OpenRouter routes (columns)
Opus 4.6
Sonnet 4.6
Haiku 4.5
Legend
  • Same pick
    Diagonal cells: both routers selected the same tier.
  • Split decision
    Off-diagonal cells: the routers selected different tiers.
  • No route
    No usable selected tier was recorded. Counts remain in the row denominator.
  • Zero prompts
    Neutral cells display 0 when no prompt has this row and column combination.

Rows: Cortex pick. Columns: selected tier. Tiers ordered by study cost; missing routes last.

Cells show prompt counts; deeper color means more prompts. Select a cell for row shares and interpretation.

  • Missing column inferred from Cortex route totals and recorded route pairs.
  • The largest split decisions exchange Opus and Sonnet: 126 prompts go from Cortex Opus to comparator Sonnet, and 83 in the other direction.
  • Of Cortex's 42 Haiku routes, 30 receive comparator Opus, 10 comparator Sonnet, and 2 have no comparator route.

Inspect individual routing calls

  • 16 selected routing excerpts; not the complete test set.
  • Original prompt lengths preserved; full prompts and model responses unavailable.
  • In this selected sample, Cortex matches 13 reference tiers and OpenRouter matches 4. These sample counts do not estimate full-cohort accuracy.

Showing 16 of 16 excerpts.

Routing calls side by side · sorted by prompt number
Routing calls side by side · sorted by prompt number
Opus 4.6Opus 4.6 · MatchSonnet 4.6 · Mis-route
Opus 4.6Opus 4.6 · MatchOpus 4.6 · Match
Haiku 4.5Haiku 4.5 · MatchOpus 4.6 · Mis-route
Sonnet 4.6Sonnet 4.6 · MatchSonnet 4.6 · Match
Opus 4.6Opus 4.6 · MatchSonnet 4.6 · Mis-route
Opus 4.6Opus 4.6 · MatchOpus 4.6 · Match
Sonnet 4.6Sonnet 4.6 · MatchSonnet 4.6 · Match
Opus 4.6Opus 4.6 · MatchSonnet 4.6 · Mis-route
Opus 4.6Opus 4.6 · MatchSonnet 4.6 · Mis-route
Haiku 4.5Haiku 4.5 · MatchSonnet 4.6 · Mis-route
Haiku 4.5Opus 4.6 · Mis-routeSonnet 4.6 · Mis-route
Opus 4.6Haiku 4.5 · Mis-routeSonnet 4.6 · Mis-route
Sonnet 4.6Sonnet 4.6 · MatchOpus 4.6 · Mis-route
Sonnet 4.6Sonnet 4.6 · MatchOpus 4.6 · Mis-route
Sonnet 4.6Sonnet 4.6 · MatchOpus 4.6 · Mis-route
Sonnet 4.6Haiku 4.5 · Mis-routeNo route · Unresolved

Prompt #0

Excerpt from a 15,919-character prompt.

...levant to your task. </system-reminder> start a worktree to solve 19 private _ functions imported from provider_endpoint_inventory │ Important │ export_scrubbed_fixture.py:29–50 │ Reviewer says follow-up PR acceptable
Prompt 0: reported route choices
Prompt 0: reported route choices
Ground truthOpus 4.6Reference label
CortexOpus 4.6Yes
OpenRouterSonnet 4.6No
  • Reported classifier probabilities: L 0.85 · M 0.14 · S 0.01.
  • L / M / S identify Opus / Sonnet / Haiku.

Methodology and limits

The archived report describes prompts from work such as coding tasks, PR reviews, agent traces, and safety-rule writing. In its reported pipeline, each prompt becomes an embedding; a lightweight classifier scores the model tiers and applies the routing rule before answer generation.

Methodology reported in the archived study
Methodology reported in the archived study
Dataset400 evaluation prompts from work traffic
Split600 training / 400 test prompts; reported overlap 0
Embedding modelOpenAI text-embedding-3-small
ClassifierRandom forest; no class weighting
Routing ruleHighest predicted model-tier score; reported threshold 0.5
EvaluationFrozen judge score matrix shared by the compared policies
  • Cortex trained on this traffic; the off-the-shelf comparator had different prior exposure, limiting generalization.
  • Original split, score matrix, and complete prompt-level routes unavailable for independent verification.
  • Session caching and production token accounting were not measured.

Archived work-memory embedding projection

Archived training-set PCA projection
Archived 3D PCA plot labeled large 169, medium 206, small 25. Its 400-prompt training caption conflicts with the 600-prompt training split.
Legend
  • Sonnet / medium
    206 prompts labeled medium in the archived image.
  • Opus / large
    169 prompts labeled large in the archived image.
  • Haiku / small
    25 prompts labeled small in the archived image.

Reported variance: PC1 13.3%, PC2 10.6%, PC3 8.2%; together 32.1%. PCA axes represent variation, not named task concepts.

The image caption describes 400 training prompts. Other archived records specify 600 training prompts and 400 test prompts. The image's cohort cannot be reconciled from the supplied records.

This projection illustrates the reported prompt distribution; it does not establish classifier performance or model-tier separability.

Routing with the cache in mind

The original report used this illustration to explain why a single-turn routing saving may disappear in a longer agent session. A reusable prefix can include system instructions, tool definitions and conversation context; its cached input price differs from the cost of processing it again.

Cache-cost illustration for one turn
0%100%
Legend
  • Warm cache hit, same model
    The archived illustration assumes a cache read costs one tenth of standard input processing. This is an assumption, not a measured result from the 400-prompt study.
  • Cold reprocess, switched model
    The illustration assumes switching models requires processing the reused prefix at the uncached input price. Actual charges depend on the provider, prefix and cache state.

These percentages apply to the reused input prefix. New input and output charges are separate; cache behavior was not measured in this study.

Share of uncached input price. Bars are ordered lowest first.

Hover, focus or select a bar for details.

  • A cache-blind router can save on a model's list rate while paying more to reprocess a long prefix. Compare the avoided cache benefit with the expected quality gain and cost of a switch.
  • The archived cache-aware routing discussion calls for tracking warm state, expiry and prefix changes; preserving a useful main-session cache; and considering cheaper subagents for narrower tasks. These production behaviors were not evaluated in this benchmark.
Cache-aware routing tradeoffs described in the original report
Cache-aware routing tradeoffs described in the original report
Short conversationsLittle reusable context accumulates; choose primarily for the required quality.
Strong model plus cheaper modelsCompare keeping the strong model's warm prefix with the quality and price benefit of moving a turn.
Cheap-only model poolWhen token-price differences are small, quality may matter more than preserving the cache.

The archive's example economics assume cache rates; actual discounts, retention and routing overhead require workload-specific measurement. The study contains no session-level token accounting.

Archived list-price illustration

Claude list rates per million tokens — June 2026 (as stated in the archived report)
Claude list rates per million tokens — June 2026 (as stated in the archived report)
Haiku 4.5claude-haiku-4-5$1.000$5.000420
Sonnet 4.6claude-sonnet-4-6$3.000$15.000147195
Opus 4.6claude-opus-4-6$5.000$25.000211198
Estimated route-weighted price over known routes
Estimated route-weighted price over known routes
Cortex400$3.845$19.225
OpenRouter393$4.008$20.038
  • Model list rates are sorted from lowest to highest input price; the route columns retain each tier's recorded counts.
  • Historical rates assume equal token volume per route; unknown routes excluded.
  • Excludes token-length variation, cache effects, batching, and router overhead.
  • List prices use a different cost basis from study units; this illustration is not a measured bill.

Current Anthropic pricing