Cortex
CortexFoundry · Datasets & Evals

Private evals on your work.Training-grade data from every outcome.

Improve your models with private evals and RL-ready data from the work your teams already do.

Every outcome becomes a training signal.

Cortex finds your highest-spend AI workflows, then captures the context, actions, corrections, and results that make them RL-ready.

Cortex · Work Unit annotation

One Work Unit. Both sides of the label.

OpenSearch retrieval fix

Preferred

Restore version-aware retrieval for accepted support answers.

Why this workflow: Support retrieval absorbs 31% of this team’s AI spend and repeats across 1,284 Work Units each month.

Models + tools

+2

Reviewers

18.4M

Tokens · 30 days

$42.6K

AI spend · 30 days

1,284

Similar Work Units

92%

Verified outcomes

Annotated automatically
ObjectiveContextModel actionTool resultHuman correctionTerminal outcome

Traceable lineage

Every source, action, review, and outcome stays attached.

Source context

01
JiraSlack

Jira issue + Slack incident

Model trajectory

02
Claude

Claude · prompt + response

Tool evidence

03
GitHubCursor

GitHub diff + Cursor test run

Human review

04

Rejected, corrected, approved

Terminal outcome

05
Shipped

Merged + retained after release

Training asset

06
RL

SFT + hard negative + reward

Rejected pathHard negative

Tune ranking weight

Claude proposed a ranking-only path that treated the symptom instead of the retrieval regression.

Evidence captured
Claude modelGitHub
  • Reviewer rejected the approach
  • Root cause remained unresolved
  • Support fixture still returned stale docs
Outcome annotationReward −1.0
Accepted pathPreferred trajectory

Fix version filtering

Claude used migration notes, corrected filter ordering, and added a regression test before merge.

Evidence captured
Claude modelGitHubCursor
  • Migration notes identified the root cause
  • Support regression fixture passed
  • PR merged and remained live after release
Outcome annotationReward +1.0

Why Cortex labeled this

3 source signals
  • Human rejected the first model path
    Claude model
  • Regression suite verified the correction
    GitHubCursor
  • Merged fix was retained after release
    GitHubSlack

Training labels from this Work Unit

4 labels · Preferred
01

SFT positive

Accepted correction to imitate

Model

Claude

Timestamp

10:42:18

Device

MacBook Pro

Surface

Cursor Composer

Application used

Cursor

Signals inferred from humans

Approved and merged

Supporting evidence found in

GitHubSlack
02

Hard negative

Rejected ranking-only approach

Model

Claude

Timestamp

10:31:04

Device

MacBook Pro

Surface

Cursor Composer

Application used

Cursor

Signals inferred from humans

Rejected in code review

Supporting evidence found in

GitHubJira
03

Regression eval

Version-filter behavior to preserve

Model

Claude

Timestamp

10:48:22

Device

CI runner

Surface

GitHub Actions

Application used

GitHub

Signals inferred from humans

Reviewer-required quality gate

Supporting evidence found in

CursorJira
04

RL-ready reward signal

Outcome reward from shipped behavior

Model

Claude

Timestamp

14:06:13

Device

Cloud runner

Surface

Release monitor

Application used

GitHub

Signals inferred from humans

Retained after release

Supporting evidence found in

SlackJira
Inspect full provenance5 events · full source trail
1contextSupport answers cite stale docs after the index migration.Issue, Slack incident, migration notes, and support fixtures
2modelDraft ranking-only path without fixing version filtering.Prompt, response, proposed diff, and missing validation
3humanReviewer rejects the path and requests root-cause evidence.Review comment and requested correction
4toolMigration notes show the version filter moved after kNN fusion.Read result, corrected diff, and support regression fixture passes
5humanMerged and retained as the accepted answer path.Outcome signal: Preferred · ship event · post-release retention

Representative private Work Unit. Customer data stays private.

One Work Unit. Five training assets.

Compile the same outcome evidence into SFT, preference, verifier, eval, and RL formats.

Cortex · Work Unit compiler

Teach the accepted path

SFT

TrainingSupervised fine-tuning

The objective, approved context, tool sequence, and shipped completion become a supervised example.

JSON training recordReady to export
{
  "source_work_unit": "WU-DEMO-4182",
  "objective": "Restore version-aware support retrieval",
  "context": ["incident thread", "migration notes", "gold fixtures"],
  "assistant_actions": ["reproduce", "patch", "test", "open_pr"],
  "accepted_completion": "Version filter applied before kNN fusion",
  "label": "preferred",
  "weight": 1.0
}

Evidence you can recompute.

Trace quality, cost, and model choice back to task-level results.

Cortex Router · ChatGPT

Tasks

500

Task-arm rows

1,500

Configurations

3

Run status

Complete

OpenAI

Luna Medium

gpt-5.6-luna · fixed

43.8%
$0.112 / task
CortexOpenAI

Router Medium

Cortex route · Luna / Terra / Sol

75.2%
$0.365 / task
OpenAI

Sol Medium

gpt-5.6-sol · fixed

79.2%
$0.501 / task

Quality is recomputed from task-level credits. Cost uses recorded input, cached-input, and output usage. Directional result; remeasure on customer traffic before production claims.

Build private evals from the work only your company can see.

Turn real work outcomes into private evals and training data you own.