Coding tasks make different demands
A small code edit and a bug spanning several files can put different demands on a model. Using one tier for both is straightforward, but the cost of that choice still needs testing.
Routing makes model selection a decision for each task. This study asks how much fixed Sol quality credit Cortex retains, and what happens to recorded model cost, when it can choose Luna, Terra or Sol.
What we compared
- Configurations. Cortex Router selects among Luna, Terra and Sol. Luna Medium and Sol Medium use fixed models, with Sol Medium as the baseline.
- Cohort. The same 500 SWE-bench Verified task IDs in each configuration: 1,500 recorded rows in total.
- Score. Recorded quality credits divided by 500 tasks, separate from the Boolean resolved flag. Resolved Luna rows carry 0.6 credit each, so its credit and resolved-task rate differ.
- Cost. Recorded model cost per task using the July 24, 2026 rate card; excludes Modal compute and router overhead.
- Limits. Totals are recomputed from the published per-task records. Those records do not include grader execution traces or establish train/test separation, uncertainty intervals, or performance on new workloads.
Overview
Performance and cost
75.2% quality credit at $0.365 per task.
Per-task distributions
What sits behind the averages.
Paired task outcomes
Where Router and Sol agree.
Recorded routing
206 tasks took a smaller model.
Published examples
Inspect a recorded task.
Methodology and limitations