Cortex
Skip to benchmark report

Cortex Router on tau2-bench

First-Turn Routing on tau2-bench

The first choice shapes every turn

A short account request and a multi-step service problem can place different demands on a model. Using one fixed model for both leaves that cost–success tradeoff untested.

Cortex OSS makes the model choice from the opening prompt and keeps it for the whole conversation. This study asks what a calibrated policy would have achieved by reusing the outcomes and costs of completed fixed-model runs.

What we compared

  • Policy and baseline. Cortex OSS selects from Qwen3.7 Plus, GLM-5.3 and Kimi K3. The fixed baseline uses GPT-5.6 Sol for every turn.
  • Paired cohort. Both configurations use the same 276 valid customer-service tasks: 247 passed for Cortex OSS and 229 for GPT-5.6 Sol.
  • Scoring and cost. Success requires a tau2 reward of 1.0. Mean model cost covers the full conversation, using historical prices as of 2026-08-31.

Overview

89.5%
Task success
83.0%
Task success
76.9%
Cortex OSS vs. Sol
+6.5 pp
Percentage points

On the same 276 paired tasks, mean model cost was $0.022/task for Cortex OSS and $0.095/task for GPT-5.6 Sol.

Performance and cost

Ten configurations

Higher and farther left is better: the counterfactual Cortex OSS point combines higher success with lower mean model cost than fixed GPT-5.6 Sol.

Task success × mean model costHigher success ↑ · Lower cost ←
Mean model cost and task success for ten configurations$0.01$0.02$0.05$0.10$0.20$0.50$1.0044.0%58.0%72.0%86.0%100.0%Mean model cost per conversation (log scale)Task success

Scroll horizontally to inspect the full chart →

Legend
  • Claude Haiku 4.5
    Fixed model · 278 valid tasks
  • Claude Opus 5
    Fixed model · 278 valid tasks
  • Claude Sonnet 5
    Fixed model · 278 valid tasks
  • Cortex OSS
    Calibrated router · 276 valid tasks
  • GLM-5.3
    Fixed model · 278 valid tasks
  • GPT-5.6 Luna
    Fixed model · 278 valid tasks
  • GPT-5.6 Sol
    Fixed model · 278 valid tasks
  • GPT-5.6 Terra
    Fixed model · 278 valid tasks
  • Kimi K3
    Fixed model · 278 valid tasks
  • Qwen3.7 Plus
    Fixed model · 278 valid tasks

Each logo marks one configuration: farther left means lower mean cost; higher means a higher share of valid tasks passed. The cost axis is logarithmic.

Cortex OSS reuses selected fixed-model outcomes and costs. Results include train and test tasks; exact per-configuration denominators appear in the table.

Hover, focus or select a logo for task counts and exact cost.

All valid tasks · highest success first · mean model cost in USD/task
All valid tasks · highest success first · mean model cost in USD/task
GLM-5.3253/27891.0%$0.032
Claude Opus 5252/27890.6%$0.809
Cortex OSS247/27689.5%$0.022
Qwen3.7 Plus234/27884.2%$0.017
GPT-5.6 Sol231/27883.1%$0.095
Kimi K3222/27879.9%$0.080
Claude Sonnet 5219/27878.8%$0.355
GPT-5.6 Terra202/27872.7%$0.043
GPT-5.6 Luna169/27860.8%$0.017
Claude Haiku 4.5150/27854.0%$0.122
Operating profile

Per-task distributions

Agent cost

USD per task
All configurations (stacked)
Tasks in bin
0$0.578$1.157
USD per task
Legend

Hover or focus a model, or hover a bar segment, to highlight it. Tap a model to pin; tap again or press Escape to clear.

2,778 valid records across all configurations. Colors identify configurations; heights count tasks.

18 unsmoothed bins with shared, rounded boundaries. The final bin keeps the full tail: 56 records above the pooled 98th percentile across all configurations.

Select a bin for exact counts, including zeros.

Exact bin counts

Scroll for all 18 bins →
Agent cost counts by configuration in USD per task; bin boundaries rounded for display. No smoothing.
Claude Haiku 4.553116594091000000000000
Claude Opus 500172017222319161819139128767
Claude Sonnet 522346483231201510171111611310
Cortex OSS27510000000000000000
GLM-5.3255230000000000000000
GPT-5.6 Luna27800000000000000000
GPT-5.6 Sol87135421031000000000000
GPT-5.6 Terra227510000000000000000
Kimi K311013726410000000000000
Qwen3.7 Plus27800000000000000000

Agent turns

Assistant turns per task
All configurations (stacked)
Tasks in bin
014.529
Assistant turns per task
Legend

Hover or focus a model, or hover a bar segment, to highlight it. Tap a model to pin; tap again or press Escape to clear.

2,778 valid records across all configurations. Colors identify configurations; heights count tasks.

18 unsmoothed bins with shared, rounded boundaries. The final bin keeps the full tail: 52 records above the pooled 98th percentile across all configurations.

Select a bin for exact counts, including zeros.

Exact bin counts

Scroll for all 18 bins →
Agent turns counts by configuration in Assistant turns per task; bin boundaries rounded for display. No smoothing.
Claude Haiku 4.50036232048283428162181910671
Claude Opus 50007231442195631122541394613
Claude Sonnet 50028228332148331721681571316
Cortex OSS0005222647363235122311164412
GLM-5.301173532523321377223742410
GPT-5.6 Luna02413362554224636712465114
GPT-5.6 Sol0211224185121444011176742414
GPT-5.6 Terra032929154928413415264109211
Kimi K30205271450224027162538116616
Qwen3.7 Plus000517224730413915221994413

Input tokens

Tokens per task
All configurations (stacked)
Tasks in bin
0157,492314,984
Tokens per task
Legend

Hover or focus a model, or hover a bar segment, to highlight it. Tap a model to pin; tap again or press Escape to clear.

2,778 valid records across all configurations. Colors identify configurations; heights count tasks.

18 unsmoothed bins with shared, rounded boundaries. The final bin keeps the full tail: 56 records above the pooled 98th percentile across all configurations.

Select a bin for exact counts, including zeros.

Exact bin counts

Scroll for all 18 bins →
Input tokens counts by configuration in Tokens per task; bin boundaries rounded for display. No smoothing.
Claude Haiku 4.5282650432423191612121411105201
Claude Opus 5011625343028221522814127410624
Claude Sonnet 503151733362416172191099551138
Cortex OSS01136553831321022386736233
GLM-5.321254633821271617364461211
GPT-5.6 Luna1034637032281795234100000
GPT-5.6 Sol7276742392720179426521102
GPT-5.6 Terra92863483728182491120010000
Kimi K3213515129222921133913534235
Qwen3.7 Plus0729414933351623576936234

Output tokens

Tokens per task
All configurations (stacked)
Tasks in bin
03,182.56,365
Tokens per task
Legend

Hover or focus a model, or hover a bar segment, to highlight it. Tap a model to pin; tap again or press Escape to clear.

2,778 valid records across all configurations. Colors identify configurations; heights count tasks.

18 unsmoothed bins with shared, rounded boundaries. The final bin keeps the full tail: 55 records above the pooled 98th percentile across all configurations.

Select a bin for exact counts, including zeros.

Exact bin counts

Scroll for all 18 bins →
Output tokens counts by configuration in Tokens per task; bin boundaries rounded for display. No smoothing.
Claude Haiku 4.53113761455538148500001000
Claude Opus 503354844404124195101110114
Claude Sonnet 50529302729322123241512984136
Cortex OSS018343022211823181114149991520
GLM-5.3026454236321620141387456112
GPT-5.6 Luna127479563211850100000000
GPT-5.6 Sol147277592819312210000000
GPT-5.6 Terra21907858206230000000000
Kimi K30931554255252015486311012
Qwen3.7 Plus0024112625292616232514131271134
Paired outcomes

Where outcomes differ

Share of all 276 paired tasks · 100% total
Legend

The four groups partition all 276 paired tasks. Segment width is the share of that same cohort; largest groups appear first.

Cortex OSS reuses outcomes from selected fixed-model runs. Results include train and test tasks; no new live router executions.

Hover, focus or select a group for exact task counts and its definition.

Routing policy

One model per conversation

Selected models across 276 valid router tasks
0276
Legend
  • Qwen3.7 Plus
    53.3% of valid router tasks.
  • GLM-5.3
    46.7% of valid router tasks.
  • Kimi K3
    0.0% of valid router tasks.

The opening prompt determines one model for every turn. Selections are calibrated counterfactuals from train and test tasks; no live router executions.

GPT and Claude are fixed comparison arms.

Selected tasks / valid router tasks. Bars are ordered highest first.

Hover, focus or select a bar for details.

Methodology and limits

Method and limits

Evaluation

  • 278-task official Sierra base (1 trial)
  • Fixed models ran live; router results are calibrated counterfactuals.
  • Task success requires a tau2 reward of 1.0.

Cohorts

  • 2 infrastructure failures excluded from their configurations.
  • Compare Cortex OSS and GPT-5.6 Sol on matching task IDs within each domain.
  • Missing outcomes and costs are never counted as zero.

Calibration

  • First-turn scores from the live Cortex sidecar feed cost-tilted gates.
  • Gates were fit only on the official training split.
  • Published results include train and test tasks; no held-out deployment validation.
  • One trial; no equal-cost comparison.

Pricing

  • Historical model prices as of 2026-08-31.
  • GPT costs use recorded token usage and historical OpenAI rates, including cached input.
  • Open models and Claude use recorded OpenRouter usage costs.
  • Latency, router overhead and total serving cost are not established.