Skip to main content

AI agents

The 20-Pound Wing Problem: Proving Agent Intent Before Write Access

Practical controls and outcomes for AI agents teams past the demo.

Headline-driven pilots skip the engineering-lead decision on ownership, proofs, and halt paths

Published
Updated
Reading time
8 min read

Key takeaways

  • Headline-driven pilots skip the engineering-lead decision on ownership, proofs, and halt paths
  • Outcome to protect: A clear build sequence the eng lead can defend
  • Prove controls under load before raising write autonomy.
  • Measure task success and incident reconstructability, not only model latency.

The demo impressed the C-suite, but the quiet cost is the engineering debt of babysitting an agent that confidently orders 20 pounds of chicken wings. You are not paying for the model; you are paying for the manual review of every hallucinated transaction. The shopper wanted two pounds of wings. The agent, interpreting "generous portion" through a degraded context window, ordered twenty. The database accepted it. The refund ticket is already open.

This is not a model failure. It is a system design failure. You built an agent that can speak, but you did not build the engineering rails that keep it from acting on its own hallucinations. The outcome you wanted was a self-correcting system that handles edge cases without human intervention. The outcome you got is a brittle system that requires a human to catch every subtle context shift before it hits the production database.

The core build decision is clear: you must build a shadow-mode validation layer before granting any write access. This is not a feature request. It is a prerequisite. You are deciding whether to prove the agent can distinguish between a user intent and a system error, or whether to accept the incident cost of finding out in production.

How do you prove intent before you allow a write?

The answer is a proof-of-work shadow run. Before any write action, the agent must simulate the transaction in a sandbox and compare the result against a known-good baseline. If the delta exceeds a threshold, the action is blocked. This is the only way to catch "confident drift" before it becomes a customer-facing bug.

The mechanism is simple but strict. The agent generates a proposed action. A separate validation service takes that proposal and runs it against a read-only copy of the state. It checks the output against the original user constraint. If the constraint was "2 lbs," and the simulation outputs "20 lbs," the delta is massive. The write is blocked. The anomaly is logged.

This control ties directly to incident cost. By catching the error in the shadow layer, you avoid the cost of a refund, the cost of a support ticket, and the cost of a post-mortem. You are trading a few milliseconds of compute for the avoidance of a high-severity incident. It is the cheapest insurance you can buy for autonomous systems.

Why does confident drift happen in long-running sessions?

Confident drift occurs when the agent treats its own previous outputs as ground truth. As the context window fills, the original user intent gets buried. The agent starts reasoning from its last action, not from the user's request. It maintains high confidence scores because the syntax is valid, even though the semantics are wrong.

This happens because the agent has no external reference point. It is a closed loop. The context window saturation causes the agent to forget the original constraint. Tool selection ambiguity leads to calling the wrong API endpoint with valid syntax. The agent thinks it is being helpful. It is actually being destructive.

The control here is not a better prompt. It is a hard limit on context window usage. You must force the agent to re-verify the original intent at every step. If the context window exceeds a certain token count, the agent must stop and ask for clarification, or it must fail safe. You cannot trust an agent that is operating on a degraded memory state.

What is the cost of a silent failure?

A silent failure is when the agent skips a step but reports success. This is the most dangerous failure mode. The user thinks the order was placed. The database shows the order was placed. But the inventory was not decremented. The payment was not captured. The state is desynchronized.

The cost of a silent failure is high because it is hard to detect. You do not get an error message. You get a discrepancy report three days later. By then, the blast radius is large. You have to trace the transaction back through the logs to find where the agent skipped the step.

The control is a state reconciliation check. After every write, the system must verify that the database state matches the agent's internal belief. If there is a mismatch, the transaction is rolled back. This is not optional. It is a fundamental requirement for any system that writes to a persistent store.

When should you halt the agent?

You halt the agent when the delta check fails. This is not a suggestion. It is a hard stop. The agent cannot proceed. It cannot retry. It cannot ask the user for confirmation. It stops. The anomaly is logged. A human reviews it.

The halt path must be owned by the engineering team. Product can define the threshold for what constitutes a "failure," but the mechanism that blocks the write must be code-reviewed and tested. If the halt path is owned by product, it will be bypassed when the business pressure is high.

The time-to-halt is critical. If the agent takes too long to halt, the user experience degrades. The agent should halt within 100 milliseconds. If it takes longer, the user thinks the system is broken. The halt path must be fast, reliable, and invisible to the user.

How do you measure the agent's reliability?

You measure reliability by the rate of human intervention. If the agent requires a human to review every action, it is not reliable. It is a chatbot with a database connection. The goal is to reduce the human intervention rate to near zero.

The metric is the percentage of actions that pass the shadow check without human review. If this number is below 95%, you are not ready for production. You are not ready to grant write access. You are still in the debugging phase.

The other metric is the mean time to detect a silent failure. This should be less than one minute. If it takes hours or days, you have a problem. The state reconciliation check must be fast enough to catch errors before they propagate.

What is the build sequence for a safe rollout?

The build sequence is Diagnose, Model, Build, Harden. First, diagnose the failure modes. What are the specific ways the agent can fail? What are the edge cases? What are the constraints?

Next, model the system. How does the agent interact with the database? What are the dependencies? What are the failure points? Create a diagram of the data flow. Identify where the shadow check fits in.

Then, build the shadow layer. This is the core of the system. It must be fast, reliable, and easy to test. Do not build the write access until the shadow layer is working.

Finally, harden the system. Add logging. Add monitoring. Add alerts. Test the halt path. Test the rollback. Test the state reconciliation. This is where most teams fail. They build the agent, but they do not build the safety rails.

Loading diagram…

How do you defer the complex parts?

You defer the complex parts by starting with a narrow use case. Do not build an agent that can do everything. Build an agent that can do one thing well. In this case, that one thing is ordering chicken wings.

Start with a read-only agent. Let it answer questions about inventory, prices, and promotions. Do not let it write to the database. This gives you time to build the shadow layer. It gives you time to understand the failure modes.

Once the shadow layer is working, you can grant write access. But you do it in stages. First, you allow the agent to write to a sandbox database. Then, you allow it to write to a test database. Finally, you allow it to write to the production database.

Each stage is gated by the rate of human intervention. If the human intervention rate is high, you do not move to the next stage. You go back to the drawing board. You fix the shadow layer. You fix the context window saturation. You fix the state desynchronization.

This is how you build a defensible system. You do not build it all at once. You build it in layers. You prove each layer before you move to the next. You do not rely on the model to be smart. You rely on the system to be safe.

The demo is a start. It is not the finish line. The finish line is a system that can handle the 20-pound wing order without a human in the loop. It is a system that can catch its own errors. It is a system that can be trusted.

This week, run a shadow simulation of your last 100 transactions. Compare the agent's output to the actual database state. Count the number of discrepancies. If the number is greater than zero, you are not ready. Fix the shadow layer. Then run it again. Do not grant write access until the discrepancy count is zero. This is the only proof that matters.

FAQ

What breaks first for shoppers find ai agents need short l?
Headline-driven pilots skip the engineering-lead decision on ownership, proofs, and halt paths That gap shows up as lost trust, longer incidents, or blocked rollouts before anyone debates model quality.
What outcome should this control model protect?
A clear build sequence the eng lead can defend. Prefer evidence operators can reconstruct over fluency in a demo.
What is a safe next check this week?
Pick one irreversible path, confirm you can halt it, reconstruct the run, and score task success in shadow before expanding autonomy.