The first choice shapes every turn
A short account request and a multi-step service problem can place different demands on a model. Using one fixed model for both leaves that cost–success tradeoff untested.
Cortex OSS makes the model choice from the opening prompt and keeps it for the whole conversation. This study asks what a calibrated policy would have achieved by reusing the outcomes and costs of completed fixed-model runs.
What we compared
- Policy and baseline. Cortex OSS selects from Qwen3.7 Plus, GLM-5.3 and Kimi K3. The fixed baseline uses GPT-5.6 Sol for every turn.
- Paired cohort. Both configurations use the same 276 valid customer-service tasks: 247 passed for Cortex OSS and 229 for GPT-5.6 Sol.
- Scoring and cost. Success requires a tau2 reward of 1.0. Mean model cost covers the full conversation, using historical prices as of 2026-08-31.
Overview
Performance and cost
Ten configurations
Operating profile
Per-task distributions
Paired outcomes
Where outcomes differ
Routing policy
One model per conversation
Methodology and limits