Discovery Is Hard. Verification Is Cheap.

by
4 min read
AI

There’s a well-known asymmetry in computer science that applies to enterprise work. Verifying that a task was completed is often easier than getting a clear picture of its components.

When I introduced the Work Unit concept, I explained how my company, Chaos Labs, identified a bunch of recurring work-related tasks spread out across teams. A key point of the article was to stress the importance of measuring and optimizing for the work being done vs. the AI prompts used to perform work.

In this article, I’ll share a framework of how we’re approaching this internally.

Work-as-Imagined vs Work-as-Done

Ask any manager to describe how work happens on their team, and you’ll typically get an “official” or much cleaner version of the process rather than a realistic picture of the actual workflow.

If you take it a step further and watch how an individual or team actually works, you’ll notice a much different process from the official documented wiki. You’ll notice some of the specifics that come with a regular workflow: the unofficial Slack threads with a group of employees, the customer-security answer accessible only to a specific team, the debugging process that exists only in one engineer’s head, and so on.

A 2024 study on procedure deviance tried to nail a number on this gap.

In this study, workers estimated their actual work departed from the written and standard work processes ~31% of the time.

This value likely understates the gap because knowledge work is typically spread across tools, changes much more quickly than companies can generally document, and is filled with important context that's implicit and impossible to reconstruct from official documentation.

The only way to infer work knowledge directly is by observing the work being done.

Process Mining for Knowledge Work

Operational processes had the same problem for decades, until process mining.

Process mining uses event logs, each tied to a case, to reconstruct how a process actually ran rather than relying on how an individual or document says it should run.

Knowledge work (including AI-mediated work) usually has no equivalent for this identifier.

Yes, the work happens, but nothing reliably marks where one task ends and the next begins.

AI activity doesn’t solve this automatically, but when captured from approved surfaces, it reveals a far clearer picture of the intent behind most work tasks.

This enables the first level of discovery: grouping related prompts, tool calls, and docs into a single Work Unit. Once grouped, Work Units can then be compared to identify recurring work types and shared characteristics.

Finding the Unknown Unknowns

In December 2024, Anthropic’s Clio system analyzed 1 million Claude conversations to identify common use cases and thousands of conversation clusters. In a separate test involving 19,476 synthetic conversations with known categories, Clio reconstructed the underlying category structure with 94% accuracy.

Together, these two tests demonstrated that Clio could discover patterns at scale and recover known category structure under controlled data.

There were some obvious limitations with Clio.

It used 2024-era models, worked only with Claude models, and didn’t have access to events before or after the model interaction. More importantly, it had no company context.

However, it demonstrated an important part of the base idea. Large volumes of AI activity can be clustered from the bottom up at useful levels of detail without requiring a human-in-the-loop to define every category in advance.

From Work Pattern to Work Policy

Once a recurring work pattern has a name and an owner, it becomes a task the company can easily manage. Cost and outcomes can now be compared across the same kind of work. Risk oversight and human-review requirements can be defined for the work they govern, and accepted outputs can easily become reusable organizational memory.

In essence, the company can define what good looks like for specific work outputs and then work backwards to set a bar for expected quality and costs.

Doing this creates a necessary feedback loop for the organization.

Future work activities can be assigned to pre-approved Work Units, and uncertain work cases are returned upstream for review. This enables a growing corpus of examples and references, allowing improvements in future rounds of discovery and growing the company’s understanding of how its work evolves.

A key policy that can be attached to validated Work Units is intelligence and resource allocation. In our semantic routing benchmark, we found that the first ~30% of frontier usage captured almost all the recoverable quality across production traffic from a subset of 40+ developers and agents.

Work Unit context allowed us to allocate intelligence based on the actual work rather than an isolated prompt. This context has become even more valuable to my team and me as the model market continues to fragment, especially with open-weight models like Kimi K3, GLM-5.2 becoming much more capable for agentic work.

The Operating Question

Every company has workflows that its employees use daily. What remains an open question is whether these workflows are visible enough to be governed.

You do not need a perfect map of the company on day one. Start with a recurring Work Unit that is driving the most spend and opportunity. Surface candidates bottom-up from actual activity, then have the teams closest to the work validate what is real, what matters, and who owns it.

Once the work becomes visible, a company can allocate the right resources and level of spend to achieve the best possible outcome, creating a win-win for the company and its stakeholders.

Risk Less.
Know More.

Get updates on our research, product, and launch.

Resources

Follow us

  • x
  • linkedin
  • youtube
Chaos LABS
Ⓒ Copyright 2026. All Rights ReservedSite monitored by Product Registry