Engineering

The Failures Your Agent Harness Catches and Never Tells You About

Agent harnesses catch errors to keep a run alive, so failed steps leave an Unset span that trace UIs render as green. Here is how to surface them.

· 14 min read
agent-harness agent-observability silent-failures opentelemetry agent-evaluation tracing
Editorial cover on a black blueprint grid reading THE FAILURES YOUR HARNESS NEVER REPORTS, with a thin line flow diagram showing a failed tool call feeding into a dashed box labelled harness except retry default, which outputs a span left unset, above a caption reading the failure stops here and two rows showing status UNSET and exceptions zero.

An agent finishes a refund workflow and reports that the refund was processed. The trace shows a completed run, no exceptions, a normal duration. Every span is green.

The refund never happened. The payment tool timed out, the harness caught the timeout and retried twice, the third attempt returned an empty body, and the output parser turned that empty body into a valid-looking object with default fields. The model read the object and wrote a confident summary.

Nothing in that chain raised. That is not a bug in the harness. It is what the harness is for.

Key takeaways

  • Agent harnesses are built to keep a run alive, so catching failures is their job, not a defect.
  • An exception that escapes a span is marked automatically, but a caught one is not, so the recovery path has to record it itself and the default status stays Unset rather than an error.
  • Five common recovery patterns hide failures: retry wrappers, step caps, stringified exceptions, coercing parsers, and quiet model fallbacks.
  • Error monitoring built on exceptions cannot see any of them, because there is no exception to see.
  • The checks that do catch them assert on the run itself: how many steps it took, which tools it called, and in what order.

Why Does an Agent Report Success After a Failed Step?

A swallowed failure is a step that failed, was caught by the harness before it could propagate, and was replaced with a default, an empty value, or a retry result. The run completes and the output is wrong.

This is worth separating from the failures agents commit on their own. Bad plans, wrong tool choices, and hallucinated answers are model failures, and we cover those in our agent failure modes taxonomy and in the piece on cascading failures in tool chains.

This post is about the layer underneath: what your agent harness does when a step fails, and why that behaviour is invisible.

Everything here happens inside one process. Once a supervisor starts dispatching work to subagents, the same problem reappears in a harder form, because a child can record its error correctly and still leave the root span clean. That case is covered in why subagent failures never roll up to the parent span.

The reason the two get confused is that they produce the same artifact. A wrong answer from a bad plan and a wrong answer from a swallowed timeout look identical in the output. They are completely different bugs, and only one of them is fixed by prompting.

Where Do Agent Harnesses Swallow Errors?

Five common recovery patterns are worth auditing first. Each is defensible in isolation. Each removes the signal you would have used to find the problem.

Retry Wrappers That Return the Last Attempt

A retry wrapper catches a failure, waits, and tries again. That is correct for a transient network error and it is why the pattern exists.

The problem is what it does when every attempt fails, and that varies more than you would expect. Some wrappers raise. LlamaIndex workflows document that “when a policy gives up, the exception propagates and fails the workflow, unless a @catch_error handler is declared”.

Others hand back the final attempt’s result, so a call that failed three times becomes an ordinary return value. Check which one yours does rather than assuming. The attempt count is usually not recorded anywhere the trace can see either way.

Step Caps That Report Partial Work as Complete

Agent loops need a ceiling or they run forever. The ceiling is the right design.

What matters is how the loop reports hitting it. If a run that exhausted its step budget returns the same shape as a run that finished its task, then “the agent stopped early” and “the agent succeeded” are indistinguishable downstream. Partial work gets promoted to a result.

Exceptions Stringified Into Model Context

A common recovery is to catch an exception, convert it to a string, and append it to the conversation so the model can react. This is genuinely useful, because a model can often route around a failed tool.

In LangChain it is also the default you get without opting in. ToolRetryMiddleware documents two options on exhaustion. The one you get without asking is 'continue', which returns “a ToolMessage with error details, allowing the LLM to handle the failure”. The alternative, 'error', re-raises and stops the run.

That default converts an error into text. The model now reasons over a stack trace as though it were data, and the run continues in a state your monitoring reads as healthy. Whether the model actually recovered is unknown to everything except the final output.

Output Parsers That Coerce Instead of Failing

Structured output helpers frequently repair malformed responses: filling missing fields with defaults, coercing types, or extracting the first object that parses. Whether strict validation is on, and what it does when it trips, is a per-library setting worth reading rather than assuming.

The result is that a response the model got wrong becomes an object that validates. Downstream code cannot tell a field the model produced from a field the parser invented.

Fallbacks That Quietly Downgrade the Model

Model fallback keeps a service up when a provider degrades. It also means the run you are looking at may not have used the model you think it used.

If the substitution is not recorded on the run, a quality drop caused by a weaker fallback model looks like a quality drop caused by your prompt.

What Does the Trace Look Like When Nothing Errored?

It looks fine. That is the entire problem, and the mechanism behind it is worth being precise about.

In OpenTelemetry, a span that stays inside a handled except block is not marked failed on your behalf.

The distinction is where the exception ends up. The Python SDK’s start_as_current_span defaults record_exception and set_status_on_exception to true, so an exception that escapes the span’s block is recorded and set to ERROR for you.

Catch it inside the block and none of that fires, because nothing escaped. The specification leaves this to language SDKs rather than mandating it, so confirm the behaviour in whichever one you run.

The trace API specification defines three status codes, and the one that matters here is the default. Unset is “the default status”. Ok means “the operation has been validated by an Application developer or Operator to have completed successfully”.

Future AGI’s tracing documentation states the two calls that change it: set_status() “sets the span’s status to OK or ERROR”, and record_exception() “attaches full exception details (type, message, stack trace) as a span event”, with the instruction to “always pair with set_status(ERROR) for complete failure context.”

Read that against a recovery path. Your harness catches the exception, so it never propagates and no automatic exception handling fires. Nothing calls record_exception. Nothing calls set_status(ERROR). The span keeps its default Unset status, which records neither success nor failure, and which trace UIs render as healthy rather than failed.

So the failure is real, the recovery is real, and the telemetry records neither. An error-rate dashboard built on span status will show a clean line through the entire incident.

The fix at the code level is small. Record before you recover, rather than instead of recovering:

from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode

tracer = trace.get_tracer(__name__)

# hides the failure: the span is never marked, the model reads "" as an answer
with tracer.start_as_current_span("call_tool") as span:
    try:
        result = call_tool(args)
    except Exception:
        result = ""

# reports it, then recovers
with tracer.start_as_current_span("call_tool") as span:
    try:
        result = call_tool(args)
    except Exception as exc:
        span.record_exception(exc)
        span.set_status(Status(StatusCode.ERROR, str(exc)))
        result = fallback_value(args)      # recovery still happens

The second version still keeps the run alive. It just refuses to lie about it.

Which Defaults Should You Change First?

These are framework-agnostic. Audit whichever harness you run against them.

Default behaviourWhat it hidesSafer setting
Retry and return the last attemptThat the call failed at all, and how oftenRecord attempt count, mark the span on final failure
Step cap returns the same shape as successPartial work presented as a finished taskReturn a distinct finish reason for cap exhaustion
Broad exception catch around tool callsThe exception type and where it came fromCatch narrowly, record, then recover
Parser repairs malformed outputWhich fields the model produced and which were filledStrict mode, so unknown or missing fields fail loudly
Silent model fallbackThat a different model served the requestRecord the model that actually responded

The pattern in that column is one idea: recovery and reporting are separate jobs. Almost every default above collapses them into one, and the reporting is the half that gets dropped.

How Do You Catch These in Evaluation Rather Than Production?

If the output is the only thing you check, a swallowed failure passes whenever the model writes a plausible summary over a bad result. Checking the output is necessary and it is not sufficient.

This is a different gap from the one where an agent passes its evals and fails in production because the rubric drifted. Here the rubric is fine, and the run it scored was never honest about what happened.

The checks that catch these assert on the run itself. Three are worth wiring first.

Step count. A run that hit its cap and a run that finished in four steps are different events. Comparing actual steps against an expected range separates them without needing to read the output at all.

Tool call correctness. A swallowed tool failure often shows up as a tool that was called and produced nothing usable, or a tool that should have been called and was not. That is checkable independently of the final answer.

Trajectory. Whether the agent executed the expected sequence catches reruns, skipped steps, and loops that a final-answer check cannot see.

None of these require you to know in advance which failure occurred. They fail on the shape of the run, which is exactly the signal a swallowed exception leaves behind.

The honest counterargument is that they may not be what catches it first.

A longitudinal study of a production agent runtime (Wu, 2026) tracked 22 incidents over eight weeks. It reports that “about 70% of silent failures were caught by human user-view observation, not tests or audits”, alongside an audit of 15 incidents finding “0% ex-ante prevention but 87% regression blocking”.

Read at face value, that is a real limit on the claim above. Assertions on run shape earn their place as regression blockers, catching a failure again once you know it exists, rather than as the thing that finds it first.

The same study records incident latencies from 13 hours to 60 days, which is the cost of relying on someone noticing.

What would falsify the case for these checks is a run where step count, tool calls and trajectory all land inside their expected ranges and the output is still wrong. That is possible, and it is the boundary of what run-shape assertions can do.

Where Does Future AGI Fit?

Everything above needs the run itself to be inspectable, which is a requirement to design for rather than add later.

Future AGI’s evals can run at three levels. When you attach one, you choose whether it scores “Spans (one step), Traces (a whole request), or Sessions (a whole conversation)”.

Scores land as a column in the trace table with a per-span result underneath. That span and trace level is where the three checks above belong.

A Future AGI trace of a research_orchestrator agent run with status OK and a duration of 2.9 seconds, showing the full span tree with plan_subtasks, research_agent, web_search, vendor_lookup, a separate vendor_lookup retry span, writer_agent, compose_brief and fact_check, with a warning badge on the root span, alongside the agent graph and the run input and output.

The run above carries a status of OK. It also contains a vendor_lookup (retry) span, which is the interesting part. The first attempt did not do its job, the harness tried again, and the run still reports success.

Whether that retry returned something usable is not answerable from the status field. Completing cleanly is not the same as being right, which is the point of scoring the span tree rather than the final message.

Three built-in evals map onto the checks in the previous section. Step Count “checks whether an agent used a reasonable number of steps to complete a task”, and the docs frame it as a way to “catch agents that loop, take shortcuts, or otherwise drift from an expected execution length”.

Tool Call Accuracy is documented for exactly this situation: “debugging agents that produce a plausible final answer but arrive at it through the wrong tool calls”.

Trajectory Match checks “whether an agent executed the right sequence of actions, not just whether it reached the right answer”.

An evaluation running directly on a Future AGI trace, scoring a ChatCompletion span from gpt-4o-mini for readability at 55.83 percent with a passing result, with the full agent trace tree on the left and the eval result tied to that specific span in the right panel.

For the failures you did not anticipate, Error Feed groups traces by failure pattern. It is scoped at “the ways agents actually fail: hallucinated outputs, tool misuse, broken workflows, safety violations, and reasoning gaps that traditional error monitoring won’t catch”.

Its scoring model separates the two signals explicitly: “a trace can score badly on one dimension without triggering a classified error, and a trace with a detected error can still score fine on unrelated dimensions”. That first combination, no classified error and a bad score, is the shape a swallowed failure takes.

Three limits are worth knowing before you rely on any of this.

The first is that these checks want bounds from you. Step Count needs expected_steps or a min_steps and max_steps range, and the docs are blunt that “at least one of these must be set, otherwise the eval fails outright”.

Trajectory Match needs an expected trajectory supplied up front and defaults to strict mode, which flags a valid alternative route that reached the same place in a different order. You do not need to predict the failure, but you do need to state what normal looks like.

The second is sampling. Eval tasks carry a sampling rate, “the share of matching rows actually scored”, and Error Feed “doesn’t analyze every trace by default”, with 10 to 20% given as a reasonable production starting point.

The release notes also record that for new tracing projects the rate “now defaults to 0%, so the Error Feed is off until you configure it”.

A swallowed failure is a low-frequency event, which is exactly the case a low sampling rate is worst at catching. Turn it up while you are hunting one.

The third is the honest ceiling. If the harness erased the signal, nothing downstream can recover the original timeout.

What these checks surface is the symptom: a step count outside its range, a tool that was called and returned nothing usable, a quality score that drops on a run your error monitoring called clean. That is enough to send you to the trace, which is where the swallowed exception gets diagnosed.

Which Silent Failures Should You Fix First?

Start with the ones that change an answer rather than delay it. A retry that eventually succeeds costs latency. A retry that returns an empty result and gets parsed into a valid object costs correctness, and nothing in your monitoring will tell you it happened.

Then separate recovery from reporting everywhere your harness catches. Record the exception and mark the span first, recover second. The run stays alive and the incident stops being invisible.

Then add one assertion on the run itself, not the answer. Step count is the cheapest place to start, because a run that hit its cap should never be indistinguishable from one that finished.

Frequently Asked Questions About Silent Agent Failures

Why Does My Agent Report Success When a Tool Call Failed?

Because the harness caught the exception before it reached your error monitoring. Agent frameworks are built to keep a run alive, so a failed tool call is commonly retried, replaced with a default, or converted into a string that goes back into the model’s context. The run finishes normally and the model treats the substituted value as a real result.

Why Does a Span Look Healthy When Something Went Wrong?

Because nothing escaped the span. The Python SDK marks a span ERROR automatically only when an exception propagates out of it, so an exception your harness catches inside the block never triggers that path. The default status is Unset, which records neither success nor failure.

Future AGI’s tracing docs describe the two calls that change it: set_status marks a span OK or ERROR, and record_exception should always be paired with set_status(ERROR). Skip both and the span stays Unset, which trace UIs render as green.

What Is the Difference Between an Agent Failure and a Swallowed Failure?

An agent failure is a mistake the model made: the wrong tool, a bad plan, a hallucinated answer. A swallowed failure is a real error in the surrounding code that the harness caught and hid. The first is a reasoning problem you fix with prompts, tools, or evals. The second is an engineering problem in your harness.

How Do You Detect Failures That Never Raise an Exception?

Assert on the trace rather than on the final answer. Check whether the agent used a reasonable number of steps, whether it called the tools the task required, and whether it followed the expected sequence. These catch runs that finished cleanly and still did the wrong thing, which exception-based monitoring cannot see.

Should Agent Frameworks Stop Catching Errors?

No. Catching errors is correct, and a harness that crashes on every transient failure is unusable. The problem is not the catch, it is the silence. Record the exception and mark the span before you recover, distinguish a completed run from one that hit a step cap, and keep substituted values distinguishable from real ones.

Frequently Asked Questions

Why does my agent report success when a tool call failed?

Because the harness caught the exception before it reached your error monitoring. Agent frameworks are built to keep a run alive, so a failed tool call is commonly retried, replaced with a default, or converted into a string that goes back into the model's context. The run finishes normally, the span is never marked, and the model treats the substituted value as a real result.

Why does a span look healthy when something went wrong?

Because nothing escaped the span. The Python SDK marks a span ERROR automatically only when an exception propagates out of it, so an exception your harness catches inside the block never triggers that path. The default status is Unset, which records neither success nor failure. Future AGI's tracing docs describe the two calls that change it: set_status marks a span OK or ERROR, and record_exception should always be paired with set_status(ERROR). Skip both and the span stays Unset, which trace UIs render as green.

What is the difference between an agent failure and a swallowed failure?

An agent failure is a mistake the model made: the wrong tool, a bad plan, a hallucinated answer. A swallowed failure is a real error in the surrounding code that the harness caught and hid. The first is a reasoning problem you fix with better prompts, tools, or evals. The second is an engineering problem you fix by changing what your harness does when a step fails.

How do you detect failures that never raise an exception?

Assert on the trace rather than on the final answer. Check whether the agent used a reasonable number of steps, whether it called the tools the task required, and whether it followed the sequence you expected. These checks catch runs that finished cleanly and still did the wrong thing, which error monitoring on exceptions alone cannot see.

Should agent frameworks stop catching errors?

No. Catching errors is correct behaviour, and a harness that crashes on every transient failure is unusable. The problem is not the catch, it is the silence. Record the exception and mark the span before you recover, distinguish a completed run from one that hit a step cap, and keep the substituted value distinguishable from a real result.
Related Articles
View all