Cortex
Skip to benchmark report

Cortex Router vs Sol

Model Routing on SWE-bench Verified

Coding tasks make different demands

A small code edit and a bug spanning several files can put different demands on a model. Using one tier for both is straightforward, but the cost of that choice still needs testing.

Routing makes model selection a decision for each task. This study asks how much fixed Sol quality credit Cortex retains, and what happens to recorded model cost, when it can choose Luna, Terra or Sol.

What we compared

  • Configurations. Cortex Router selects among Luna, Terra and Sol. Luna Medium and Sol Medium use fixed models, with Sol Medium as the baseline.
  • Cohort. The same 500 SWE-bench Verified task IDs in each configuration: 1,500 recorded rows in total.
  • Score. Recorded quality credits divided by 500 tasks, separate from the Boolean resolved flag. Resolved Luna rows carry 0.6 credit each, so its credit and resolved-task rate differ.
  • Cost. Recorded model cost per task using the July 24, 2026 rate card; excludes Modal compute and router overhead.
  • Limits. Totals are recomputed from the published per-task records. Those records do not include grader execution traces or establish train/test separation, uncertainty intervals, or performance on new workloads.

Overview

75.2%
Quality credit per task
$0.365
USD per task
27.2%
Compared with fixed Sol
94.9%
Relative quality credit
  • Same 500 SWE-bench Verified task IDs across all three configurations
  • Luna Medium and Sol Medium: fixed models; Cortex Router: selects Luna, Terra, or Sol
  • Fixed baseline: Sol Medium

Read upward for more quality credit and leftward for lower model cost. Cortex records 27.2% lower model cost than fixed Sol, with 4.0 percentage points lower quality credit.

Model cost and qualityThree measured points · log cost scale · higher and farther left is better

Quality credit per task

Mean model cost per task (USD, log scale) · $0.08–$0.60

Legend

Higher is more quality credit; farther left is lower model cost. Quality starts at 40%; cost uses a log scale. Only these three configurations were measured.

$68.25 lower model cost across 500 tasks: Cortex $182.31 vs fixed Sol $250.56. July 24, 2026 rates exclude Modal compute and router overhead.

Quality credit differs from resolved-task rate. Luna: 365 resolved × 0.6 = 219 credits / 500 tasks = 43.8%.

These points do not establish an equal-cost frontier, uncertainty intervals or performance on new workloads. Select a logo or legend row for exact values.

Performance and cost

75.2% quality credit at $0.365 per task.

Read these charts together: higher quality credit and lower model cost are better. Cortex retains 94.9% of fixed Sol’s quality credit at 72.8% of its model cost.

  • Per-task values = recorded quality credit or model cost total / 500 tasks
  • Historical rate card: July 24, 2026 (2026-07-24)
Quality credit per task · higher is better
0%100%
Legend
  • Sol Medium
    Sol Medium: one fixed model for every task
  • Cortex Router
    Cortex Router selects Luna, Terra, or Sol per task
  • Luna Medium
    Luna Medium: one fixed model for every task

Quality credit / task. Bars are ordered highest first.

Hover, focus or select a bar for details.

Mean model cost per task (USD) · lower is better
$0.00$0.55
Legend
  • Luna Medium
    Luna Medium: one fixed model for every task
  • Cortex Router
    Cortex Router selects Luna, Terra, or Sol per task
  • Sol Medium
    Sol Medium: one fixed model for every task

Model cost / task. Bars are ordered lowest first.

Hover, focus or select a bar for details.

Quality and cost totals · 500 unique tasks per configuration
Quality and cost totals · 500 unique tasks per configuration
Sol Medium50039679.2%396$250.56$0.501100.0%100.0%
Cortex Router50037675.2%376$182.31$0.36572.8%94.9%
Luna Medium50021943.8%365$56.22$0.11222.4%55.3%
  • Cortex vs fixed Sol: 27.2% lower model cost; 4.0 percentage points lower quality credit
  • No evidence for total production serving cost, latency, or savings after router overhead
Per-task distributions

What sits behind the averages.

All three configurations appear in every chart, with 500 task records each. The curves preserve the study’s independently peak-normalized shapes. Its smoothing procedure is not supplied; curve heights compare shape, not task counts or probability mass.

  • Median and p90 are recomputed from every recorded task, using linear interpolation between ordered observations
  • Original display ranges are retained; counts above the displayed range are listed below each chart and remain in the statistics
  • Effective cost includes cached-input pricing; the source defines agent steps as assistant messages, command executions, and file changes
Effective model cost

Archived source shapes · each curve peaks at 1

10

Effective model cost (USD)

Legend

Colors identify configurations. Each archived curve peaks at 1, so height compares shape, not task counts or probability mass; the smoothing method is unavailable.

Median (50th percentile) and p90 (90th percentile) include all 500 tasks and the tails beyond this range. Rows run lowest median first.

Hover or focus a curve or legend row to highlight it and inspect exact statistics.

Agent steps

Archived source shapes · each curve peaks at 1

10

Agent steps (steps)

Legend

Colors identify configurations. Each archived curve peaks at 1, so height compares shape, not task counts or probability mass; the smoothing method is unavailable.

Median (50th percentile) and p90 (90th percentile) include all 500 tasks and the tails beyond this range. Rows run highest median first.

Hover or focus a curve or legend row to highlight it and inspect exact statistics.

Input tokens

Archived source shapes · each curve peaks at 1

10

Input tokens (tokens)

Legend

Colors identify configurations. Each archived curve peaks at 1, so height compares shape, not task counts or probability mass; the smoothing method is unavailable.

Median (50th percentile) and p90 (90th percentile) include all 500 tasks and the tails beyond this range. Rows run highest median first.

Hover or focus a curve or legend row to highlight it and inspect exact statistics.

Output tokens

Archived source shapes · each curve peaks at 1

10

Output tokens (tokens)

Legend

Colors identify configurations. Each archived curve peaks at 1, so height compares shape, not task counts or probability mass; the smoothing method is unavailable.

Median (50th percentile) and p90 (90th percentile) include all 500 tasks and the tails beyond this range. Rows run highest median first.

Hover or focus a curve or legend row to highlight it and inspect exact statistics.

Paired task outcomes

Where Router and Sol agree.

Compare the recorded Boolean resolved flags on the same 500 task IDs. The shared bar partitions the cohort into both, Router only, Sol only, and neither.

Router and fixed Sol · same 500 tasks
0%100% · 500 tasks
Legend

Width is the share of 500 paired tasks, largest first. All four outcomes are shown.

Outcomes compare resolved flags, independently of quality credit. They do not explain why a routing decision succeeded or failed.

Recorded routing

206 tasks took a smaller model.

  • Luna or Terra: 206 tasks (41.2%)
  • Sol: 294 tasks (58.8%)
How the same 500 tasks were assigned
0 tasks · 0%500 tasks · 100%
Sol MediumFixed model for all 500 tasks
Legend
  • GPT-5.6 Sol
    294 Router tasks · 58.8% of the cohort
  • GPT-5.6 Luna
    112 Router tasks · 22.4% of the cohort
  • GPT-5.6 Terra
    94 Router tasks · 18.8% of the cohort

Bars show each model’s share of 500 Router tasks, largest first. Every task has one route.

Fixed Sol uses one model for all 500 tasks. Route counts alone do not establish optimal choices.

Published examples

Inspect a recorded task.

  • Choose any of the 500 task IDs to compare all three configurations
  • Scores and usage only; no prompt text or answer excerpts

Load the examples to compare results for every task.

Methodology and limitations

How the study was run.

Scope and quality

Each configuration uses the same 500 SWE-bench Verified tasks: 1,500 task results across three configurations. Quality is total recorded credit divided by the task count, separate from the Boolean resolved flag. Resolved Luna tasks receive 0.6 credit.

Generation and grading

Source-reported method: the study describes the Codex harness running in official SWE-bench task images on Modal at Medium reasoning. Fixed arms requested Luna or Sol; Cortex selected a model before agent execution and kept that choice fixed for each task.

The report describes 500 completed generations and grades per configuration, with patches graded using the canonical SWE-bench test specification. It reports retrying infrastructure failures instead of counting them as unresolved model outcomes. These procedures cannot be independently verified from the recorded scores and usage: execution traces, patches and retry logs are unavailable.

Historical model cost

Mean cost is total recorded model cost divided by 500 tasks. Costs use the historical rates below for uncached input, cached input and output tokens; they exclude Modal compute and router overhead.

Historical rate card · 2026-07-24 · USD per one million tokens
Historical rate card · 2026-07-24 · USD per one million tokens
GPT-5.6 Luna$1.00$0.10$6.00
GPT-5.6 Terra$2.50$0.25$15.00
GPT-5.6 Sol$5.00$0.50$30.00

Cached input is priced once at its discounted rate; reasoning output is not counted again. Rates apply to each task’s recorded token totals.

Recorded cache use · cached input tokens / all input tokens
Recorded cache use · cached input tokens / all input tokens
Luna Medium92.2%
Cortex Router91.6%
Sol Medium91.4%

Cache shares use total cached-input tokens divided by total input tokens, rather than averaging task percentages. Re-measure cache behavior and provider pricing on your workload before applying these comparisons.

Routing and interpretation

The study measures one Cortex Router operating point. It does not establish a threshold sweep, train/test separation, uncertainty intervals, or performance on new workloads. The paired results show where outcomes differ, but do not establish which routing decisions caused the quality gap. Task examples expose recorded scores and usage, without prompts or answers.