← Field Notes

2026-09-28 · Trends

Agents Inherit the System They Land In

Enterprise AI agent deployments fail for reasons that sit below the model. The causes reported in 2026 — no machine-readable definition of done, no access to the actual work, evaluation coverage that drifts — are all properties of the system the agent was dropped into. An agent inherits the substrate it lands in.

The 2026 numbers point at the floor, not the ceiling

Agents are now standard equipment. Gartner puts the share of enterprise applications shipped or updated in Q1 2026 carrying at least one embedded agent at roughly 80%, against 33% in 2024, and expects about 40% to include task-specific agents by year end, up from under 5% a year ago.

Deploying is not the same as working. Forrester's root-cause analysis, as reported this year, finds 22% of agent deployments sitting at negative ROI twelve months in, and attributes those failures to unclear success criteria (41%), insufficient tool or data access (33%), and drift in evaluation coverage (26%). IDC and Microsoft put the return for deployments that do reach production scale at 171%. The distance between those two numbers is the entire subject.

IBM's Institute for Business Value surveyed 2,000 C-level technology executives across 33 geographies and 19 industries between January and April 2026. Two-thirds reported being accountable for AI systems they do not fully control; 70% said teams deploy faster than IT can track; 11% said they were ready for the scale of agent deployment they expect within the year.

These are analyst and vendor survey figures, gathered on different bases, and they disagree on magnitude — read them for shape rather than precision. The shape is consistent, and it is the part worth keeping:

Not one of the named failure causes is a model-quality problem. Every one of them describes the system the agent was deployed into.

Three things an agent needs that most systems don't have

A definition of done that a program can read. In most businesses, "finished" lives in someone's judgment or in a document written for a human reader. An agent handed that has to infer a terminal state, and inference is how you get a system that reports success because a loop ended. MagenticOS wrote up a live instance in What "Completed" Means: a task returned Completed with the model provider unreachable, because an empty result set is a valid result set and nothing in the path was entitled to object. "Unclear success criteria" is that problem wearing a business-case costume.

Reach into the work, not into documents about the work. A third of the negative-ROI deployments are attributed to insufficient tool or data access, and that is rarely a permissions ticket. The work itself — the delivery, the filing, the matter, the obligation — exists as PDFs, spreadsheet tabs, and group texts, so the only thing an agent can be given access to is a reading list. Access is a data-model property before it is a security setting.

A surface stable enough to be evaluated against. Evaluation coverage drifts because the thing being evaluated moves. When a rule lives in a prompt or a runbook, changing it changes behaviour everywhere at once and silently invalidates every test written against the old version. When the rule is a constraint in the schema, it moves in one place and the system tells you when something violates it. The related argument at the OS layer is capability-based security: authority granted per action and typed, rather than per app and assumed.

Sprawl is the symptom, not the disease

The IBM finding that two-thirds of technology leaders are accountable for systems they don't control reads as a governance story. It is more useful as an architecture story. Every function bought its own agent; nothing consolidated underneath them. The result is an organization with more automation and less coherence than it had two years ago — an expensive place to arrive.

The instinct at that point is to buy an orchestration layer — an agent to manage the agents — which adds a tier above the problem. What was missing was never coordination; it was a shared model of the work that more than one system agrees on. Which is the same claim we made when defining an operating system for an industry, from the other direction: dashboards, copilots, and now agents all sit on a record layer, and none can be better than it is.

What the substrate looks like when it exists

In LegiOS, an obligation is not a paragraph — it carries its citation, jurisdiction, effective date, and deadline as fields, which is what lets the compliance calendar be re-derived rather than hand-maintained when a date moves, the argument in an earlier post. Its scope is bounded: it states when a question exceeds what software should decide and routes the person to an attorney. That is a definition of done that includes "not mine" — the hardest one to write, and the one that keeps a system honest.

LeafIQ puts routes, orders, deliveries, and AR on one data model, so the delivery that gets invoiced is the same delivery the compliance trail describes, rather than a reconciliation between a driver's texts and an accounting system. In Casebound, a matter cannot open without a cleared conflicts check, a draft from the Lexwright engine cannot reach a client without a recorded attorney disposition, and every mutation is audited, including from an admin account — rules in the schema rather than in a policy binder. LeoLog takes it to the limit: each entry is hashed with SHA-256 and encrypted on the device before anything leaves it, then anchored on-chain via Base and Bitcoin through OpenTimestamps — the content never touches the chain — so a proof bundle can be checked by anyone at verify.leolog.io without a LeoLog account.

None of these was built to make agent deployments succeed, and we are not claiming an agent-readiness feature we haven't shipped. They were built because regulated work needs a record that survives a skeptical reader — the requirement that showed up last week as a pricing problem. Agent-readiness is a side effect of that discipline, which is the actual claim: the substrate is not an AI feature, and you cannot buy it from the vendor who sold you the agent.

Three questions before the pilot

Ask them about the specific workflow, not about the company:

Where is "done" written, in a form a program can read? If the answer is a person's judgment or a document, the agent will invent one.

Can the agent reach the work, or only artifacts describing the work? If the underlying event was never recorded as data, no amount of tool access reaches it.

When the rule changes, what changes? A prompt, a configuration value, or a constraint. Only the third gives you something to evaluate against next quarter.

If the honest answer to all three is "someone knows," the pilot is not an agent problem waiting on a better model. It is a substrate problem, and it will still be there after the vendor bake-off. The same discipline makes a research process worth following — rules fixed before the decision rather than narrated after it, the standing argument on the MacroSavant Substack — and it is why we build the record layer first.

We publish here as we learn. Subscribe by RSS, or write to hello@mutagenic.io.