When a thunk has been running for a while and you want to know how it is doing — or you have been asked to "clean it up" — there are several places to look: the work items, the Errors, Token Costs and Time Analysis tabs in Monitor, and the design review. This article tells you which to look at first, and why. Monitor is available to thunk admins, and running a design review needs the Owner or Admin role on the thunk.
This article has two parts. Concepts explains how to decide what matters most. Details walks through each check in order: where to look, what to look for, and which article to read next.
Concepts
Check what does the most harm first
Problems in a running thunk differ in how much harm they do while nobody notices them. Check them in that order:
Stuck work — work items that have stopped and will not move on their own. Nothing alerts you, and the work is lost until someone acts.
Errors — runs that failed. They are recorded where you can see them, but each one is work that did not finish as designed.
Wrong results — work items that finished but did the wrong thing. Nothing fails, so these are found only by looking at the results.
Design problems — the causes behind the first three, which a design review can point to.
Time — how long the work takes.
Cost — what the work spends on AI.
Time and cost are worth improving, but only once the thunk is doing the right thing: making a thunk that gives wrong results faster or cheaper does not help anyone. Keep notes as you go — a finding in one check often explains another, and the design review in particular tends to explain what you saw in the first three.
Not every failure is an error
An error is a run that the platform could not complete: a tool failed, a model refused, a connection was missing. Errors are recorded and listed on the Errors tab.
Many workflows also have business outcomes that are failures for your process but not errors for the platform: a request handed off to a person, a lookup that found nothing, an output left empty. The step finished normally, so nothing appears on the Errors tab. These are found in the work items' own data, which is why checking results is a separate step from checking errors.
Look at recent work, and at a sample of it
A thunk's history can hold far more work items than you can read. Look at a recent window — the last few days — so that what you find reflects the thunk as it is now, not as it was before its last change. Count how the outcomes fall across that window, then open a few work items of each kind, especially the unusual ones.
Details
1. Stuck work
Start here, because stuck work does not fix itself.
Is the whole thunk paused? A paused thunk does no work at all. Check Monitor → Pause / Restart. See Pause and restart a thunk.
Which work items are waiting? In the Work Items list, filter the AI Task column to Waiting. The filter also includes Waiting for tool. A work item that is Waiting needs a person — an answer, a decision, or Finish step; one that is Waiting for tool is waiting on a tool, most often a reply on a conversation. Waiting is often the design working as intended; it is a problem when nobody is expected to act, or when the work items have been waiting for days. See What Working, Waiting, and Waiting for tool mean.
Why did a step stop? Open a waiting step and read how its agent ended the run. An agent that pauses for a person when its instructions meant it to finish is an instruction problem, and every work item that takes the same path will stop in the same place. See How an AI agent ends a run.
Work items that never finish. Filter the list on Status to the steps before the end of the workflow, and look for work items that were created long ago. Find out why they stopped before you close them. If stalled work should be closed automatically, set Auto-finish inactive work items after (hours) — it moves them to Final, it does not do the remaining work; see Data retention.
2. Errors
Open Monitor → Errors. The error-rate chart at the top shows how many runs errored each day; the list below it holds the most recent errors, newest first, each with its category, message and remediation note. See Analysis of AI behavior for everything on the tab.
Read the trend before the list. A single error is often a passing failure in an outside system. A rising error rate, or a rate that jumped on the day the thunk changed, is a problem.
Group the errors before reading them one by one. The same error repeated across many work items is one problem with one fix.
Read the category first. It tells you whether the fix is yours (the instructions or tools), an administrator's (a connection or credential), the platform's, or whoever supplied the input. See Troubleshooting Errors.
Open the failed run. Expand it from the error, and use What went wrong? for an explanation of that run. The error you see is often not where the failure started — see Reading a failed run.
3. Wrong results
Check what the work items actually produced.
Count the outcomes. Filter the Work Items list on the properties that record how each work item ended — a status, a result, a "needs a person" flag — and compare the counts. A share of hand-offs or empty results that is higher than you expect is the finding, even when no single work item looks wrong. To keep a view you will come back to, save it as a report.
Look for values that should not be there. The text "null" or "-" in a field, a list where one value was expected, or an output that disagrees with the property it came from usually means the instructions do not say clearly what to write.
Look for properties nothing fills in. A property that is always empty is either not needed or not mentioned in any step's instructions.
Open a few of each outcome. Read the steps of a typical success, a typical failure, and anything unusual. When an agent did the wrong thing, find the step and the turn where it went wrong. See AI reliability troubleshooting and Analysis of AI behavior.
4. Design review
Run a design review. It checks the thunk's properties, steps and tools for problems that commonly lead to the failures above — a property no step uses, a property whose description does not match its type, a step given inputs its instructions never refer to, a tool whose arguments are not constrained. Read its findings alongside what you found in checks 1–3: a finding about a step you already saw going wrong is the one to fix first. See Workflow Review.
When you fix what you found, turn a few of the work items that went wrong into automated tests, so the fix stays fixed when the thunk changes again.
5. Time
Open Monitor → Time Analysis. Find the step that accounts for most of the time, then the work items that were slowest. Separate time you can change (how the step's instructions lead the agent to work) from time you cannot (waiting on a person, the platform's own time). See Time Analysis.
6. Cost
Open Monitor → Token Costs. Find the step that spends the most per work item and why: a large tool result read on every turn, many turns, or a larger model than the step needs. See Token Costs.
