This week you can send reports on a schedule, add AI checks to a step's finish condition, and find any inbound message in a filterable table. Why so long? now also tells the platform's time apart from your automation's.
New
Send reports on a schedule. Every thunk now has a report automation beside its reports in Monitor → Reporting: an AI agent that acts on your saved reports. Write what you want in its Settings — "every Monday at 8am, email the Overdue Items report to [email protected] as an Excel file" — turn on a schedule, and it runs on its own. It can send a link to a report's live page, attach the report as a CSV or Excel file, or write a short summary, and you can ask it for something right now in its Execution chat. Add a Query Filter to the schedule and it runs only when a work item matches, so a quiet day costs nothing. See Send reports on a schedule with the report automation.
Finish conditions can include AI checks. A step's Step Finish Condition can now hold statements an AI model must judge true before the workflow proceeds, such as The Receipt is an expense receipt, not some other document, plus a one-switch check that the step's output values match their property descriptions and types. AI checks can read the content of file and image properties, so a wrong document is caught even when the file itself is present. If a check does not pass, the step stays open for a person, just as it does for an unmet condition. See Control when a step runs and proceeds.
Find any inbound message. Each channel in Inbound Requests now lists all of its messages in a table — status, when each was received, subject, sender and an excerpt — instead of the newest few with a Load more button. Filter the table to find, for example, every message that failed to route, and share the filtered view by its link. A channel's settings, such as Instructions and API Keys, are now in the bar at the bottom of its tab. See Inbound Requests.
AI coding agents can unit-test a custom tool. Through the Builder API (including from an AI coding agent over MCP), you can now give a custom tool test cases, run them, and read the graded results, without running a work item. Each case is an input plus what must be true of the tool's result. Running all of a tool's cases saves the results, which then appear in the tool's Testing section and in Workflow Review, the same as tests run from the tool builder. The tool runs for real, so a tool that calls another system or sends mail does that once per case. See Testing custom tools.
The conversational API tells you when a conversation has ended. The
GET /{convoId}/eventsstream now sends aconversationEndedevent when the workflow behind a conversation finishes or is ended early, so a chat client can stop waiting without polling. ItsendReasoniscompletedwhen the workflow ran to its end andendedwhen it was ended early.GET /{convoId}/inforeturns the sameendReasononceendedis true. See Conversational interface via a REST API.
Improvements
Why so long? explains more of a run. A run's breakdown now shows Queued (waiting for the platform to pick the run up), Setting up (preparing the step before the agent starts) and, for a step, Wrapping up (saving the results after a run ends) as bands of their own, so Between runs covers only time when nothing was running or waiting to run. Monitor → Time Analysis also counts queued and setup time toward a run's working time, so figures for new runs are slightly higher than before. The breakdown marks a tool call repeated with exactly the same inputs, and the explanation reads the instructions and tools the step ran with, so its advice is about how the step is written. Why so long? and What went wrong? now run on the thunk's planning model. See Analysis of AI behavior.
AI model changes. Fast now works on Gemini models served through Vertex AI: a thunk whose Execution service tier is Fast uses Google's priority processing, which responds faster and costs more, and falls back to Standard if priority capacity is short. Gemini models reached through the Gemini API still run at Standard. Kimi K2 Thinking now shows (deprecating soon) and is no longer offered when adding a model to an environment. GLM-5.2 has been withdrawn by its provider and can no longer run; GLM-5.3 is available as an experimental model. See Supported AI models and what happens when they change.
Test model checks streaming too. In Model Definitions, Test model now sends its request both normally and streamed. If the streamed request fails, the result warns that workflows with Use streaming LLM API turned on will fail with that model. A new model, or one whose provider, model or API base you change, is now tested as soon as you save it. See Orgs, Roles, and Environments: Enterprise Administration Guide.
Deleted work items are kept for 7 days. For new thunks, Data retention period for deleted items (days) now defaults to 7 days instead of 0, so a deleted work item stays in the trash for a week before it is permanently removed. Existing thunks keep their current setting. See Data Retention.
Builder API and Thunk API refinements. Reading a thunk's design with
get_definitionnow returns a count of each connection's reference targets instead of every target, which could overrun a coding agent's reading limit; ask for the targets withrefPositions. Work-item queries can return the newest work items first withorder: "desc"and alimit. A filter or returned-field name the thunk does not have is now a clear error that lists the valid properties, rather than an empty result. See Reuse thunk interfaces and plug in connections.
Fixes
Forms enforce required fields and show what the agent knows. A form's Submit or Save button now stays off until every required field is filled, including required fields inside each row of a list or inside a group of fields. When the agent asks someone to review or correct details it already has, it now fills them into the form instead of opening it empty.
Quick successive changes to tools and properties all stick. Turning tools on or off in quick succession — in a step's AI Tools, in Connections, in Custom-built tools, or in an MCP Export Interface — or quickly adding, removing or binding input and output properties could lose one of the changes. Each change now builds on the ones before it.
No more limit of ten. An inbound message routed to more than ten work items now reaches all of them instead of failing, and can no longer be routed to a work item in a different thunk. If you have more than ten OAuth credentials, your AI tools can now use all of them, not ten chosen arbitrarily.
Automated tests grade JSON data and finish large re-grades. A test rule on a JSON Object property, or a list of objects, now compares the data itself rather than its formatted text: equals ignores field order, and contains passes when one item has those fields. Re-grade test assertions on a large test suite no longer starts over when it runs out of time; work items it did not reach stay pending evaluation, and the new Re-grade ungraded button on the Test Results tab finishes them. See Automated tests for workflow steps.
Model settings name the model that actually runs. In a thunk's AI Settings, a model setting left at Default now shows the thunk owner's default for that setting, and the planning, security-checking and custom-tool-builder models each show their own default. When model settings are not set, the owner sees Fix it, which saves their current defaults on the thunk so everyone can see which models it uses.
