Skip to main content

Improve a live thunk safely

Clean up or upgrade a thunk that does live work: sort fixes from the owner's decisions, test in a copy with stand-in tools and real inputs, and keep a list of what to undo before going live.

A thunk that has been doing live work for a while collects problems: work items that get stuck, values that come out wrong, instructions written before a better platform feature existed. Fixing them is risky, because the thunk acts on real systems — it sends email, updates records, files tickets — and often belongs to someone else. This article describes how to improve such a thunk without touching the original or anything it writes to, and how to prove the improvement before it goes live.

This article has two parts. Concepts explains the approach. Details walks through it in order. Start with a health check, which tells you what is worth fixing: see Checking a thunk's health: where to start.

Concepts

Keep what it decides, fix how it runs

Set the goal before you start: the improved thunk makes the same business decisions as the original, and stops failing while it carries them out. Who gets handed to a person, what counts as a match, when to try again — those are the owner's rules, even where you would choose differently. A cleanup that quietly changes them is a different workflow, and nobody agreed to it.

Two kinds of finding

Every problem you find is one of two kinds:

  • A fix. The thunk does not do what its design intends: work items stuck on a step, values that are invented or garbled, a success recorded when nothing happened, a property bound to a step that never fills it in. You can fix these.

  • A decision for the owner. The right behavior is a business choice: try again or hand off, accept a near match or flag it, handle a request that names two people. Only the owner can answer these, so they become questions, not changes.

Sort every finding into one of the two before you change anything, and agree the list of fixes with the owner.

Never test against the real systems

Make every change in a copy of the thunk, never in the original. In the copy, replace each tool that changes something outside the thunk with a stand-in: a tool with the same name that returns what the real one would have returned, and records what it would have done. Then the copy can run real inputs end to end, and nothing reaches a customer, a record or an inbox.

Details

1. Ask the owner what only they can answer

Go through this list and write a question for each line that applies to the thunk. Most apply to any thunk that acts on an outside system. If you skip a line, note why.

  • Failures. When an outside system is briefly unavailable or busy, should the thunk try again — how often, and how far apart — or hand the work to a person at once?

  • Near matches. Should a nickname, a former name, a typo or a different format count as a match, or be flagged for a person?

  • Ambiguous requests. When one request names several records or people, or two sources disagree (a form field and the message text), should the thunk act on all of them, on the one that agrees, or hand off?

  • Requests it was not built for. A link to another site, an id in another format: translate it, look it up, or hand off?

  • "Already done." When the outside system answers that the action was already taken, is that a success, and should the work item still close?

  • Repeats. When the same request arrives again, should the thunk skip it, do it again, or tell someone?

  • Who may ask. May anyone request this action on any record, or must the requester be checked first?

  • Replies. Should the requester get a reply, should only an internal note be left, or both — and in what words?

  • What arrives. What exactly does the trigger — an email, a form, a webhook — send, and which fields are always there? A fix that relies on a field that is sometimes missing will fail on those work items.

  • Accounts. Whose account does each connection use — a person's or a shared service account — and who replaces its credentials when they expire?

  • A safe place to try. Is there a test system, a sandbox or a test record where the real action can be tried without consequences?

  • Retention. How long may work items that hold personal details be kept? See Data retention.

  • The description. Where the thunk's description promises something no step does, should the step be built, or the promise removed?

Write each question so the owner can answer it in a line: the question and its options, what the thunk does today, the work items that show it, and your recommendation if you have one. Include any work item where the thunk may already have done the wrong thing in another system, so the owner can check it there.

2. Make a copy to work in

Copy the thunk and work only in the copy. Before you run anything, read Copying a thunk to try out changes: the copy keeps the original's phase, runs with its owner's connections, and can still reach every real system the original reaches.

Pause the copy until its live tools are replaced, and turn off any tool library a step does not need — a step with every tool turned on can still reach a real system through a tool you forgot. See Pause and restart a thunk and AI tool libraries.

3. Replace the live tools with stand-ins

For each tool that changes something outside the thunk — sends email, updates a record, uploads a file, replies to a ticket — and for each lookup whose answers you need to control, add a custom code tool that stands in for it:

  • Give it the same name as the tool it replaces, so the steps' instructions keep working unchanged. Remove the live tool from the copy first, or turn off the tool library it comes from.

  • Answer with what the real system said. A custom tool's Call history, and the tool calls in a step's run, show what the real tool was given and what it returned. Use those recorded answers, looked up by the input — this ticket number returns that ticket — so the copy sees exactly what the live thunk saw, including the odd values. See Analysis of AI behavior.

  • Fill the gaps from the work items. Where there is no recorded answer, build one from the original work items, including the awkward cases you found: a record with an empty field, a file that looks right but is not.

  • Be stricter than the real system. Have the stand-in reject the bad requests the real system quietly accepted — a missing value, text with stray formatting, a placeholder link — so a test fails where production stored junk.

  • Fail on purpose. Choose a few inputs that make the stand-in fail the way the real system can: unavailable, too busy, "already done". Then the thunk's failure paths get tested too.

  • Record what it would have done. A stand-in for an action returns what it would have sent or changed, so a test can check "replied with this text, tagged it that way".

  • Return the same ids every time for the same input, so a test can expect an exact value.

Run each stand-in on its own before a step uses it, including an input it should reject. See Automated tests for custom tools.

Recorded answers hold real customer data. Keep them inside the thunk, where the thunk's access controls apply, and do not paste them into documents or messages.

If several copies need the same stand-ins, build them once in a separate thunk and share them as a library. See Modular reuse of tools.

4. Build tests from real work items

Create test work items in the copy from the inputs of real work items in the original — one for each behavior, not one for each work item:

  • the usual case that works;

  • each kind of failure the health check found, named so you can refer to it later;

  • each failure the stand-ins can produce on purpose; and

  • the cases where the thunk should do nothing: a message that is not a request, a missing input, a request it should hand off.

Expect exact values where you can — the ids the stand-ins return, the allowed values of a property, an empty field where nothing should be sent. Keep AI-graded assertions for generated text. Tests check the work item's own values, not the state of the outside systems, which change on their own. See Building a test plan.

Run one test work item first to check that files open and the stand-ins are connected, then run the rest.

5. Make the agreed fixes

Make only the fixes the owner agreed to, and leave the decisions for the owner's answers. Common fixes:

  • A way out of every failure. Every failure or refusal ends with the work item marked as failed and needing a person, never stuck on a step; the last step always finishes; success is recorded only from what the action actually returned.

  • Values the agent copies. Pass short keys instead of long values, check values at the property and at the tool, and guard any action you cannot undo. See AI reliability troubleshooting.

  • Properties. Make an input that is often missing optional, so nothing is invented to fill it; give results with a known set of values a fixed list of allowed values; fix wrong types; remove bindings to properties that do not exist.

  • A gate before risky actions. A property set by the first step that classifies the request, a condition on the later steps, and a human approval for the cases that should stop.

  • Instructions. An objective and a numbered process, what to do when a value is missing, empty values left empty rather than written as "null" or "-", and each decision made by one step. See How to Write Effective AI Instructions.

  • Platform features the thunk predates. Allowed values on properties instead of format rules in the instructions, Auto-finish inactive work items after (hours) so stalled work does not pile up, a better-suited model for a step.

Write down each change and why you made it as you go. The write-up at the end is built from these notes.

6. Run the tests, and compare models

Run the whole set and read the results. For each failure, decide which of four it is: the thunk is wrong (fix it and run again), the test is wrong (correct it and use Re-grade test assertions), the behavior waits on an owner decision (leave it failing and point at the question), or the copy cannot reach something it needs (say so). Do not weaken a test to make it pass.

Once the results are stable, you can run the same set on one or two other AI models and compare pass rate, time and cost. See Move a thunk to a new AI model. Finish with a design review of the copy.

7. Keep a list of what to undo before going live

From the moment you make the copy, keep a list of every change that exists only to test it safely — and that must be undone before the copy, or its fixes, can run against the real systems:

  • remove the stand-ins and restore the live tools;

  • turn back on the tool libraries you turned off for testing, if production needs them;

  • remove any test-only text from the instructions;

  • turn off test-only settings, and set the production model; and

  • run the tests and the design review again with the live tools in place, then try a few real work items under watch before the thunk takes its full load.

Add to the list as you make each change, not from memory at the end. It is what stands between "the tests pass" and "this is safe to run for real".

8. Hand over to the owner

Give the owner a short write-up:

  • The result in a sentence, and the test pass rate.

  • The changes, each with why it was made, and with the test-only parts marked.

  • The original failures now fixed, each with how often it happened in the samples you looked at, an example work item, and the change and tests that cover it.

  • What is left: the list of what to undo before going live, known limits of the fixes, credentials to replace, data to clean up in other systems.

  • The questions for the owner from step 1, numbered.

Add the owner to the copy as an admin, if they do not already own it, so they can see the work. See Users and access controls.

Did this answer your question?