Full transcript
What Would Make You Stop the AI Run?
EP048 · Oct 5, 2026 · 17m 10s
Source record
Sources cited in this episode
Public references used to prepare the companion episode.
AI Change Desk | EP048: What Would Make You Stop the AI Run? Runtime 17:10 Transcript of the approved final master, including the reusable disclosure and song voice lines. RSS.com transcription compared against the approved rendering inputs; minor recognition and spelling errors corrected.
You hit stop. The screen goes quiet. Then your phone lights up. The supplier got the email. Here is the part nobody wants to explain. You stopped the task you could see. Another task was already doing the sending. And tomorrow's scheduled run is still waiting for breakfast. Imagine that happening with a supplier shortlist. A document changed. An approval expired. You caught it, and you acted. So why is the workflow still moving? This is a fictional example, not a reported incident.
But the control question is very real. Apparently, stop needs a job description now. Today, we are going to make that description useful. What would make you interrupt an AI run? What exactly would you expect to stop? And what evidence would you need before starting it again? A stop request is something you send. A stopped workflow is something you verify. Those are not the same event. Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow.
Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization. One additional disclosure as the desk evolves. My current role includes privacy work. so you will hear me pay closer attention to purpose, access, retention, deletion, and accountability. I will not discuss nonpublic work here. These are my personal views, and they do not represent the state of Oregon or any other organization.
This is AI Change Desk. Thank you for choosing to tune in. Welcome back to AI Change Desk. I am Michael. This is our Monday episode for October 5th. I checked the public sources on Sunday, October 4th. That matters because today's story includes a new launch, living documentation, and a preview with changing capabilities. I will keep those boundaries clear. In episode 47, we asked who checks a safety claim and who responds when the reviewer finds something important. Now take the next step.
Somebody has decided the work should pause. What has to happen for that decision to become true? This is not another episode about naming an incident owner. It is about the awkward space between giving an instruction and finding out what the connected systems actually did. On September 29th, OpenAI introduced DOTS, Agents Designed to Work Over Time, using connected applications. Its announcement describes a rollout, not universal availability. It also distinguishes current offerings from specialist enterprise pilots. The practical change is work that keeps going after the person who requested it has turned elsewhere.
That can be useful. I would very much like software to handle the tedious part without requiring me to become its full -time chaperone. But when work continues across tasks and applications, our operating instructions have to describe more than a conversation. The new question is not whether a chat stopped answering. It is whether the work stopped doing. We have one main change and two supporting signals today. different kinds of stop controls, the difference between redirecting work and canceling it, and evidence that helps you decide when intervention is warranted.
Then we will work through a small exercise. No live accounts, no real customer data, and no heroic experiment with the production stop button. The useful detail in OpenAI's controls documentation is the distinction among three operations. Pausing a dot stops its current main task. It does not stop every delegated task or cancel future scheduled runs. Delegated work and recurring schedules have separate controls. The documentation also says stopping work does not undo completed actions. That is the documented boundary for this product.
It is not a claim that every agent platform behaves identically. Now return to our fictional supplier exercise. The main task is comparing proposals. A second task is preparing supplier messages. A recurring job will check for updated proposals tomorrow. Those are three pieces of work. There could also be an action already accepted by another service. and there could be a completed message. The person pressing stop may mean all of that. The control they pressed may cover only one piece.
That is an interface problem worth understanding before an emergency. It is also a responsibility problem. Somebody needs to know where the rest of the work lives. I would ask the workflow owner to draw the work, not just the application architecture. Here is the active task. Here is the delegated task. Here is the future schedule. Here is the outside service holding an accepted request. Here is the action that already happened. For each one, name the control and its owner.
Then, name the evidence that would confirm its state. Do not assume disabling access and canceling work are interchangeable. The right sequence depends on the system and the situation. An authorized administrator needs to understand that sequence, including any effect on investigation and recovery. This is where I would slow the room down. If the screen says paused, what is the noun? The conversation? The worker? The schedule? The entire workflow? A label without a scope leaves the operator guessing. And the operator is usually the person who gets the phone call when guessing becomes visible.
The takeaway is modest but important. Before approving autonomous work, ask for its stop map. Not a promise that someone can take over. A map of the work, the controls, and their limits. The second signal comes from OpenAI's October 2nd guide to building with the GPT -6 family. It describes steering a running response through the developer interface. The dedicated documentation makes the timing distinction explicit. Acceptance means the new input is queued. It does not mean the model has acted on it.
Steering does not cancel tools that already started or reverse earlier actions. This is an API behavior. Do not treat it as a universal statement about every button in every application. Here is why it matters to the person running the workflow. Suppose the supplier instructions say, prepare the approved messages. Then the authorized person changes the scope. Hold the messages. The shortlist needs another review. There are at least three moments to distinguish. The correction was submitted. The correction was accepted.
The relevant work actually changed. Only the last one answers the operational question. And even then, you still need to account for earlier effects. This is not about distrusting every acknowledgement. An acknowledgement can be accurate and useful. It just may be answering a smaller question than yours. Your request has been received. Excellent. So has my request for fewer status meetings. Receipt is a start, not an outcome. For a high -consequence workflow, I would record both. The request identifier and its time.
Then the observed change, its time, and the evidence. If you cannot establish the effect, say unknown. Do not translate unknown into cancelled. Do not translate unknown into safe to retry. The outside system may have accepted the first action, even if your application never received the response. Sending another request could create a second action. That is a general distributed workflow risk, not a claim that it happened in either product discussed here. Ask the service owner how results are reconciled and how the workflow prevents duplicate effects.
Test that answer with disposable data before relying on it. The organizing principle is simple. Keep the instruction, its acknowledgement, and its effect separate. That distinction becomes more valuable as work becomes more autonomous. Of course, you cannot interrupt every time an agent thinks, or waits, or takes longer than the person watching would prefer. So what should trigger the stop decision? Microsoft's October 2nd article about insights in Foundry discusses using traces to identify recurring agent behavior. One illustration concerns repeated planning and a proposed termination improvement.
The proposal still needs evaluation. Microsoft's documentation labels insights as a public preview. It calls for reviewing the evidence and validating proposed changes. It also warns that insights are not guaranteed to be real-time. A finding can help you investigate. It is not, by itself, an emergency brake. That distinction gives us a better way to think about stop conditions in our own exercise. Start with what would make the work unacceptable to continue. The agent reaches outside the approved purpose.
The required approval is no longer valid. It repeats an action without making meaningful progress. or the evidence needed to supervise the work disappears. Those are proposed conditions to adapt and test. They are not universal thresholds that I can set for you. An expensive run might still be doing useful work. A quiet run might be correctly waiting for a reviewer. A fast run might be executing the wrong thing beautifully. So compare the concerning behavior with a healthy example.
What does useful progress look like in this workflow? What does a legitimate wait look like? What would count as completion? In the supplier exercise, a healthy wait might mean the messages exist as drafts and cannot be sent until an authorized reviewer makes a decision. An unhealthy loop might mean the same check runs repeatedly, with no new information and no route to an owner. The difference is not how busy the interface looks. It is whether the next action can improve the outcome within the purpose and authority already approved.
That is the question I would give the team before I ask them to tune a timer. Then test both sides of the proposed intervention. Can it interrupt the behavior you want to prevent? Can it avoid interrupting the healthy comparison? Otherwise, you can replace one operational problem with another. The agent no longer gets stuck. It also no longer finishes anything. Very low incident volume, extremely low usefulness. The decision is not simply whether a warning looks convincing. It is whether the evidence supports a change and whether that change works under the conditions you care about.
Let us finish the supplier scenario. You pause the main work. The delegated task owner confirms that task has stopped. The scheduling owner confirms tomorrow's recurrence is disabled. You still have one sent message and one outside service request with an unknown outcome. This is not the moment for a cheerful green checkmark. The sent message needs a business decision. Was the recipient right? Was the content approved? Does it require a correction or escalation? The unknown request needs an investigation owner.
Someone must reconcile it with the receiving system before deciding whether another attempt is appropriate. And the people waiting on the process need an update. Not an assurance that everything is fine, an accurate statement about what is known and what happens next. For our exercise, that might sound like this. We have paused the review workflow and disabled its next run. One supplier message was sent before the pause. We are checking a second request with the service owner. No additional messages are authorized while that check is open.
The process owner will provide an update this afternoon. That is clearer than saying the AI is off. It gives the next person something they can act on. There is a privacy reason to be precise here too. Stopping new work does not answer every question about information already copied, shared, or retained. You may need to find drafts, exports, and downstream copies. You may also need to preserve relevant evidence. Use the organization's incident privacy, and records procedures to decide what to restrict, retain, correct, or remove.
Do not improvise deletion instructions from a podcast, and do not put sensitive material into a new tool just because the old workflow is under review. For today's exercise, use invented records and test destinations. There is no reason a rehearsal needs a real person's information. This is a deliberate connection to our earlier privacy episodes. Purpose, access, and retention still matter after the agent stops. The new decision is what the organization does with work that has already crossed a boundary.
An interruption is not a completed recovery. Give each unresolved effect an owner and a next check. Keep it open until the evidence supports a disposition. Now comes the tempting part. The error is fixed. The owner is back online. Somebody says, can we just resume? Maybe, but first, what exactly would resume? In our example, the supplier shortlist changed during the pause. The person who approved the first message is no longer the person authorized to approve the next one.
If the workflow simply continues from its old context, it could faithfully finish work nobody wants anymore. That is why I would treat restart as a fresh decision, not the administrative opposite of stop. Confirm the objective still makes sense. Confirm the data and destinations are still appropriate. Confirm who can approve the next consequential action. and reconcile the work that already happened. Then choose how to proceed. Resume the unfinished portion. Start a revised task with a narrower scope. Keep the workflow held while a dependency is unresolved.
Or close it because the work is no longer needed. Those are different business decisions. They deserve different records. Yesterday's approval is not a gym membership. It should not renew forever because nobody found the cancellation page. I would also give the restarted work a short observation window. Watch the behavior that caused the interruption. Check whether the revised path creates a different problem. Keep the accountable person reachable. This does not mean every restart needs a committee. Proportion the decision to the consequence.
A disposable test has a different risk from a workflow sending messages or changing records. But somebody should be able to explain why this run is allowed to continue now. Not why the technology was approved six months ago. Why this run, with this purpose, These inputs and this remaining uncertainty is authorized now. Here is the practical work for this week. Give one workflow 45 minutes. Use a tabletop first with synthetic records. If you are listening while driving, leave the notes for when you are safely parked.
The worksheet will carry the steps. Bring the process owner and someone who understands the controls. Include the relevant privacy or security person when the example touches their responsibilities. For the first 10 minutes, draw five cards. Main task. delegated work, future schedule, completed action, unknown result. Place each card next to the person who could establish its current state. If no one owns a card, that is a finding. Do not fix it by putting everyone's name on it. For the next 10 minutes, choose a trigger.
An approval expires during the run, or the same action repeats without useful progress. Pick one. Not every possible disaster at once. Write what should stop, what may safely continue, and what must wait for a human decision. Include a healthy waiting example so the team does not confuse patience with failure. Spend the next 10 minutes walking through the response. Who makes the request? Who has the authority and access to act? What acknowledgement would they receive? What separate evidence would establish the effect?
Give the team one awkward fact. The main task is paused, but a result is unknown. Ask whether anyone is about to retry it. If so, what evidence makes that safe? For the next 10 minutes, decide about restart. Change one condition while the workflow is held. Use a revised recipient, purpose, or approving person. Can the team notice that change? Can it explain which earlier approval no longer covers the work? Can it separate already completed actions from unfinished ones?
Use the last five minutes to write a disposition. What stayed held? What was reconciled? What may restart under whose authority? And who owns the unresolved item and its deadline? The output is not a thick policy document. It is one page someone can use under pressure. Record the trigger and the scope. Record the controls and their owners. Keep acknowledgements separate from evidence of effect. List completed and unknown outcomes. Then record the restart decision and remaining limits. Mark any technical claim you could not test.
A good tabletop can reveal the missing test. It cannot substitute for running that test safely. If the workflow needs a live technical rehearsal later, get the right authorization and use a controlled environment. Do not turn this assignment into an unannounced production interruption. One final thought. People should be able to interrupt work without pretending they already understand every consequence. The system should help them discover those consequences, not reward the first person who declares the issue closed. Useful autonomy is not the absence of human involvement.
It is knowing when involvement matters and making that intervention specific enough to work. This week, do not just ask whether your agent has a stop button. Ask what would make you press it. Ask which work it controls. And decide what evidence you would need before restarting. The next time someone says, we stopped the AI, ask them to finish the sentence. Which work stopped? What still needs attention? Who verified it? That is a conversation worth having before the supplier calls.
I am Michael. This is AI Change Desk. Thanks for listening. I will see you next week. This is AI Change Desk. See you next time.