This week Time Analysis and Token Costs chart every work item over time, the hosted Chat UI can take form answers and follows a person across devices, and plug-in connections show and let you change the interface they require. Coding agents can also run a Design Review and grade test work items through the Builder API.
New
Run a Design Review and grade test work items from the Builder API. After you author a thunk over HTTP or MCP, you can start the same Design Review as Deploy → Review (
run_review), then pollget_reviewfor a pass/fail verdict and the findings to fix. A coding agent can also create a test work item together with its expected results (create_work_itemwithtests), change or remove those assertions, and regrade the current results (batch_work_item_tests), then read each assertion's pass or fail fromget_work_item_state. Exact expectations are graded as rules; plain-English statements are graded by AI. See Workflow Review and Automated tests for workflow steps.Time Analysis and Token Costs plot your work items over time. Monitor → Time Analysis and Monitor → Token Costs now draw every work item in the scope as a dot at the time it started, with each day's median, 90th and 99th percentile running through them, so a workflow that is getting slower or more expensive shows as a rising cloud. Dots can be coloured by Cache hit rate or by LLM model. The daily chart has a new Show control: Total per day (as before), Median work item per day, and Runs per day — set beside the total, the median tells more work apart from slower or pricier work. Time views no longer show a Mean, because a few slow items pull the average above anything typical. See Time Analysis: how long your thunk takes and Token Costs: what your thunk spends on AI.
The hosted Chat UI takes form answers and follows the person. With Form responses on for a conversation that uses the Thunk Chat UI channel, the agent can attach a form to its message and the person fills it in right there in the chat; their answers become their reply. The sidebar now lists the conversations a person started on any device they sign in from, instead of only the ones started in that browser; use End conversation to close one. Integrations can do the same through the conversational REST API: read
responseFormon a message, submit toPOST /{convoId}/message/{messageId}/formResponse, and list a user's conversations withGET /conversations. See Set up the hosted Chat UI for your thunk.See, export, and change the interface a plug-in connection requires. A plug-in connection's page now has collapsible Interface, Binding and Tools sections and opens on its tools. View interface shows every tool's description and inputs and exports the interface to a file; Change interface… replaces it from a file, first listing the tools it adds, removes, and changes. Interfaces now match on tool names, input names and input types, so a provider that rewords a description still matches. A property can also take its type from a connected tool: give it the Reference type and pick a position in a tool's result, such as a plug-in connection's
get_account. See Reuse thunk interfaces and plug-in connections.Choose when the AI saves a step's results. By default the AI writes results to the work item as it works — after each search, for example. A new When the AI saves results setting can instead have it save once, just before it finishes or pauses the step, which takes fewer AI turns and leaves no half-finished values in the grid. The trade-off: results not yet saved are lost if the step times out or fails. Set the thunk-wide default in the thunk's AI settings under Reliability Guardian, and override it for a single step in that step's AI Settings. See AI reliability mechanisms.
Improvements
Why so long? separates waiting from working. The breakdown now opens with the Wall-clock time and the Agent working time side by side. Time a run spent waiting on a person is its own figure and never counts as tool time or Platform overhead. Where the time went has a new In order view that lists every call as it happened; opening a call shows where it falls among similar calls, such as Call 3 of 6 in its group, and each model turn shows its token counts and output speed. See Analysis of AI behavior.
Custom code tools can wait on a person and message a conversation. Give a code tool a Conversation input and, when it runs on a work item, the script can send on that conversation, wait for the person's reply, and keep going — including when one code tool calls another. A code tool can also pass its conversation to another custom tool. After a run, Try it and Call history show a Console block and a separate Tools called block. See How to build a tool that uses a conversation.
Choose how many requests a custom tool runs at once. Every custom-built tool can now take several requests in one call, including API, database and classifier tools. While Allow batch processing? is on, a new Requests run at a time setting (1 to 10) decides how many run together. Code tools start at 10, AI tools at 3, and all other tools at 1. Each request gets its own result, so one failing request no longer fails the others. See AI Tools.
Upload files through a REST API connection. A REST API connection now includes a
generic_api_post_binarytool for upload endpoints such as ServiceNow attachments. Give it the file's MIME type and a link the file can be downloaded from, and it POSTs the file's contents with that MIME type. Like the PUT, PATCH, and DELETE tools, it is off until you enable it on the connection. See Connecting to External Systems and Tools.Test and manage model definitions more easily. Each card in the org admin Model Definitions tab has a Test model button that sends one short request to that definition and shows how long it took and what it cost, flagging a $0 cost as unpriced usage. Definitions list all their settings as labelled rows, Assign Model Definition always shows the current list and has a Filter box, and admins whose role comes through a group now get the management controls. GPT-4.1, GPT-4.1 mini and Gemini 3.0 Flash now show (deprecating soon); thunks using them keep working. See Orgs, Roles, and Environments: Enterprise Administration Guide.
Watch a custom AI or custom code tool's activity while it is still running. When a workflow step calls a custom AI or custom code tool that waits — for example for a conversation reply — the calling step's activity log now shows that tool's own log live, not only after the tool finishes. See Custom Tools.
Fixes
Replies reach an agent that is waiting for them. An emailed reply, a reply through a webhook, a form answered from an email link, and a form submitted in the hosted Chat UI now resume the step that was waiting, instead of leaving it waiting. On a conversation that sends several assistant messages in one turn,
agentWorkingnow stays true until it is your turn.Quick Connect can finish signing in to an MCP server. Adding an MCP connection with Quick Connect used to fail after you allowed access, with Token exchange failed: Missing required parameters, even when the same server signed in fine from Claude Desktop. That authorization now completes.
Thunks set to a retired model keep running. A thunk whose model had been retired failed to run. It now runs on a stable model in the same family and says so in the chat. A run's If this loop had run on cost comparison also no longer offers the model the run already used, which showed a saving that wasn't there.
Custom tools and plug-in connections do what their controls say. A custom tool's Enabled switch is now what decides whether the agent gets it — if one showed Enabled but was missing, turn it off and back on. Calling an imported thunk's workflow tool creates a work item again, and Remove works on a plug-in connection whose name has capital letters or spaces.
Google Chat is no longer a conversation channel. Conversation properties no longer offer Google Chat. A leftover column that still uses it will not send; change it to Email or Thunk Chat UI. The Google Chat webhook tool is unchanged. See Conversation Properties for Email.
