Skip to main content

Reading a failed run

Understand why a run of your thunk failed: how failures are shaped, who can fix them, and what to do next.

When a step fails, the error message tells you what was reported. This article is about the next question: why it happened, and who can do something about it. Concepts explains what a run is made of and how a failure travels through it. Details explains how to read the shape of a failure and decide what to do.

For the meaning of a specific error message and its standard fix, see Troubleshooting Errors.

Concepts

What a run is made of

Each time an AI agent works on a step, it makes a run. A run is a sequence of model turns, in which the agent reads the instructions and the work item's data and decides what to do, and tool calls, in which it acts: searching, reading a file, writing to a spreadsheet, sending an email. Each tool call has a status and an output.

Some tools are themselves AI agents. A custom AI tool runs its own nested run, with its own instructions and its own tool calls. When something fails inside it, the failure is handed back to the run that called it.

Where it broke is not always where it shows

The error you see is the one at the outside: the step that was running when the failure reached the top. The cause can be deeper — a tool inside a custom AI tool, two or three levels down. Finding the run the failure started in is the first step to understanding it, because that is where the instructions and tools that need attention live.

The platform's classification

Every error is shown with a category (LLM, Platform, Thunk Definition, Connection/Tool, User Error and so on) and standard advice for that kind of error. That classification comes from the platform's catalog of known errors and is the authoritative answer to "what kind of error is this?". Reading a failed run adds the part the catalog cannot know: what this particular run was trying to do when it hit the error.

Who can fix it

A failure is fixed by one of four people, and the fixes are different jobs:

Who

What they change

You, the author

The instructions of a step or a custom AI tool, which tools a step can use, the thunk's properties

An administrator

Connections, credentials, permissions and the environment the thunk runs in

Thunk

A defect or a limit in the platform itself — contact support

Whoever supplied the input

The file, link, address or value the run was given

Knowing who can act comes before knowing what to do: advice to edit instructions is no help to someone waiting on a connection they cannot change.

Details

Find the run where it started

Follow the failure inward. If the outer error is a custom AI tool that failed, open that tool's run and look at what failed there. Repeat until you reach a run whose own tool call, or the agent's own decision, caused the failure.

Read the shape of the failure

Most failures fall into a handful of shapes, and the shape says more than the message does:

  • The same tool failed several times with the same output. This is a hard block: a missing permission, a record that does not exist, a configuration that is wrong. It will not fix itself, and running the step again will not help.

  • The same tool failed several times with different outputs. The agent was trying variations of something that cannot work. Usually the instructions ask for something the tool does not offer.

  • A tool started and never returned. This is a timeout, not a rejection. The external system was slow or the work was too large; the fix is different from a refusal.

  • A person denied an approval. This is not a fault. The run asked a person to approve something, and they said no.

  • No tool failed at all. The agent itself concluded that it could not go on. Its last message before stopping is the best evidence of why — read it before suspecting anything else.

Compare the instructions with the tools

Most failures an author can fix are a mismatch between what the instructions ask for and what the step's tools can actually do: an instruction to update a record through a tool that can only read, to search a source no enabled tool can reach, or to produce a value no input supplies. Read the wording that sent the agent to the failing tool, and check it against that tool's description. See How to Write Effective AI Instructions for writing instructions that match the step's tools.

Decide whether to try again

Run the step again only when the evidence points to a passing problem: the AI provider timed out, was rate-limiting, or returned a server error. For anything else — the same output on every attempt, a permission, a missing record, a mismatch with the instructions — running it again repeats the failure.

Be honest about what the record cannot say

The run records which tools were called, what they returned, and what the agent said. It does not always say why an external system returned what it did. When it does not, the useful next step is to check that system directly (the record, the folder, the permission) rather than to guess.

Use What went wrong?

On a failed run, the What went wrong? button beside the error explains the failure in these terms. It is also available on failures listed under Monitor → Errors. The explanation reads in this order:

  1. what broke, and what kind of failure it was;

  2. who can fix it (for example, You can fix this or An administrator needs to fix this);

  3. why it happened;

  4. the steps to take, with links to the help articles that cover them; and

  5. the tool calls behind it, starting with the run where it broke.

When the same failure repeated exactly, the explanation says that retrying will not help. When the cause cannot be established from the record, it says so and shows how confident it is, rather than presenting a guess as fact.

Did this answer your question?