Full transcript
Work Agent Receipt Check
EP037 · Jul 20, 2026 · 19m 54s
Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.
Before we get into this week's change, I owe you a quick update. The Desk went quiet for a little while. I did not explain that very well. I left GTO, and I now work for the State of Oregon.
At the same time, my family and I started planning a move to another state. That means a new job, a new home, kids, dogs, logistics, and the kind of spreadsheet that eventually develops its own weather system.
So I took some time to get the important things organized. The show is continuing. But the operating model is changing. AI Change Desk will now be one focused episode each week. Short enough for a commute. Long enough to explain one emerging change, what it means for your organization, and what you should do next.
The former midweek episode is going away. Those signals will now be combined into the weekly show instead of becoming a second release just because the calendar says Wednesday. From time to time, I may also bring in an assistive guest created through NotebookLM.
When that happens, I will disclose it, bound it to the source material, review it, and remain responsible for the final editorial judgment. The tools can assist. They do not get the last word. As always, these are my opinions.
They do not represent the State of Oregon, GTO, or any other organization. Thank you for giving me the room to reset. Now, back to the Desk. Imagine an AI agent finishes a task at two thirteen in the morning.
It found the files. It checked the context. It called three tools. It changed a record. It sent the follow-up. And then it left one final message. Done. Great. Done according to whom? Which version of the file did it use?
What did the tools change? What did the run cost? Who reviewed the result? And if the answer is wrong, where is the evidence that lets another person reconstruct what happened? Because done is not evidence. Done is a status word with excellent public relations.
It is a green check mark wearing a blazer. When AI starts completing work instead of merely drafting text, the organization needs more than a completion message. It needs a receipt. If the agent did the work, the system should be able to prove the work.
Intro music
Welcome back to AI Change Desk. I am Michael. Today is Monday, July twentieth, twenty twenty-six. And this is episode thirty-seven: Work Agent Receipt Check. As I am source-checking this episode today, the freshest operating signal is not another benchmark race.
It is a measurement question. OpenAI published a scorecard for the AI age on July seventeenth. The post argues that organizations should look beyond seats, active users, licenses, and cost per token. Instead, it proposes four questions. How much useful work gets done?
What does a successful task actually cost? How often does AI get the work right? And does each AI dollar produce more value as usage grows? That is vendor-authored framing. It is not an independent return-on-investment study. But the operating direction is useful.
Because most AI dashboards are still very good at telling you that something happened. Tokens were used. Requests were made. People logged in. The graph went up and to the right, which remains the preferred migration pattern for graphs seeking executive sponsorship.
The graph does not know whether the work was useful. The scorecard starts with one workflow and one definition of done. For support, done might mean the customer issue was resolved. For engineering, it might mean a code change passed its tests.
For legal, the OpenAI post says it could mean a contract was reviewed accurately and on time. The important phrase is not contract. It is quality bar. The task is not successful because the agent stopped running. The task is successful because the result met a defined standard in the system where the work actually matters.
Completion is an event. Acceptance is a decision. That distinction sounds small. It is not. If an agent drafts a customer response, completion means the draft exists. Acceptance means the response is accurate, appropriate, sent through the approved channel, and connected to the right customer record.
If an agent changes code, completion means the patch exists. Acceptance means the tests passed, the review happened, the deployment boundary held, and the rollback path is real. If an agent prepares a forecast, completion means there is a spreadsheet.
Acceptance means the inputs were current, the formulas survived, the assumptions are visible, and finance is not discovering a mystery tab during the meeting. This is why every delegated workflow needs an outcome owner. Not merely an agent owner.
Not merely an admin owner. Someone has to own the sentence: This result is acceptable for this purpose. If nobody can say that sentence, the agent is not finishing work. It is producing material for somebody else to finish later.
That may still be valuable. But it is a different operating claim. The second question is cost. And this is where token dashboards become a little too pleased with themselves. OpenAI's scorecard says the full business cost of a successful task includes employee time, human review, retries, and rework.
That is the right direction. The cheapest model call can still produce the most expensive outcome if the team has to rerun it four times, repair the output, re-check the sources, and apologize to somebody named Karen in procurement.
The receipt therefore needs two numbers. What did the system consume? And what did the successful outcome cost? Those are not the same number. The first may include model usage, tool calls, compute, external services, and elapsed runtime.
The second adds review time, corrections, retries, rework, and any downstream failure the team had to absorb. If the task failed, do not quietly remove it from the denominator and call the pilot efficient. That is not measurement.
That is a tiny accounting costume. The receipt should show whether the result was ready to use, needed correction, or needed escalation. Those are also the three dependability outcomes named in the OpenAI scorecard. Ready to use. Needs correction.
Needs escalation. Simple categories are useful because they force the organization to count the human work hiding behind the machine's confidence. And they create a better management question. Are we reducing effort while quality holds? Or are we moving the effort into review, exception handling, and cleanup where the AI dashboard cannot see it?
Hidden review is still work. Hidden rework is still cost. The third part of the receipt is authority. The same OpenAI post says dependability requires clear boundaries before AI moves from drafting to taking action. What data can the system access?
What systems can it use or change? When should a person review or approve an action? That is the bridge from a scorecard to governance. An outcome can be correct and still be operationally unacceptable. The right invoice can be sent to the wrong customer.
The right file can be placed in the wrong workspace. The right code can be committed to the wrong branch. The right answer can come from data the agent was never supposed to touch. So the receipt cannot only say what happened.
It has to say what the agent was allowed to do. Which identity ran the task? Which data sources were available? Which tools were called? Which records changed? Which steps required approval? Which exception path was available? And who had stop power while the work was still in motion?
This is where episode thirty's standing-permission question meets episode thirty-three's runtime-budget question. Episode thirty-five asked what approval surface survives after AI leaves the desk. Episode thirty-six asked who gets frontier capability and under which gate. The receipt is where those controls meet.
It is the small artifact that says: Here is the outcome. Here is the authority. Here is the cost. Here is the evidence. Here is the human decision. The second signal makes that evidence problem more concrete. OpenAI introduced ChatGPT Work on July ninth as an agent for longer, more involved tasks.
According to the official release notes, Work can research and analyze information, operate across connected apps and files, create finished deliverables, and run scheduled tasks. Users can follow progress, answer questions, change direction, and approve important actions. Then, on July sixteenth, OpenAI updated the desktop experience.
Chat and Work conversations appear together in Recents. Projects are available in the desktop app. Cloud Work conversations can continue across web, mobile, and desktop. Local conversations stay on the computer. Codex remains a separate view. That is a product detail with an organizational consequence.
Continuity is not custody. A conversation that follows you across devices may look like the same work. But a cloud-backed Work thread and a local conversation do not create the same evidence path. They do not necessarily have the same storage behavior, administrative visibility, or review surface.
And the release note does not answer every retention, training, residency, legal-hold, or compliance question. Do not invent those answers from a feature announcement. Same interface does not mean same control boundary. For a Work task, the receipt should record the execution context.
Was the work local or cloud-backed? Which Project supplied context? Which connected apps or files were used? Which tool actions occurred? Which approvals were requested? Which deliverable was accepted? And where can a reviewer find the evidence without opening fourteen tabs and one screenshot named final, final, actual final, two?
File naming remains undefeated. The Microsoft signal shows why the receipt also needs control-plane state. Microsoft documentation says that, effective July first, certain security capabilities for Copilot Studio and Microsoft Foundry agents require an eligible Microsoft Agent three sixty-five license.
For licensed tenants, Microsoft says observability logs and the agent registry become central evidence surfaces. The agent inventory in Advanced Hunting is moving from one table to another. Some threat-detection alerts move to Agent three sixty-five observability logs.
Third-party agent discovery moves toward registry sync. And there is one very specific warning operators should notice. For tenants using existing Agent three sixty-five real-time protection rules set to block, Microsoft says those rules stop blocking on July first unless they are redefined in the new policy experience.
That statement is scoped. It does not mean every Microsoft control stopped working. The documentation specifically distinguishes Agent three sixty-five real-time protection from Copilot Studio protection through Defender for Cloud Apps. Licensing and tenant state matter. But the broader lesson travels well.
A control receipt needs a version. The organization may have approved the workflow in June. The license may change in July. The inventory source may move. The alert path may change. The blocking rule may need to be recreated.
The agent may still run while the evidence surface underneath it has changed. If your receipt only says security approved, it becomes stale the moment the control plane moves. The receipt should say which policy was active, which registry held the inventory, which log source captured the event, and when those facts were last verified.
Microsoft's Copilot Studio what's-new page, last updated July sixteenth, reinforces the same problem from the builder side. The page lists organizational-data access through Microsoft IQ. Reusable skills. Persistent memory. A Windows cloud P C connection through M C P.
Delegation to other agents. And multiple primary-model choices. Some capabilities are generally available. Some are preview. Availability can depend on product, region, configuration, and licensing. The point is not that every organization has every feature. The point is that the receipt has to capture the actual configuration of the agent that ran.
Not the product brochure. Not the demo environment. Not what the team remembers approving three months ago. The actual agent. Its actual memory state. Its actual skills. Its actual model. Its actual connections. Its actual delegated work. Before the action block, it helps to say what a receipt is not.
It is not a transcript dump. It is not forty pages of tool output placed in a folder nobody can search. It is not a screenshot of a green check mark. And it is not the agent writing a paragraph about how carefully the agent performed the task.
Self-assessment is useful. It is not exactly a hostile witness. A receipt is a compact index into the evidence. It tells a reviewer what outcome was accepted, where the source material came from, what actions occurred, what approval was required, what the successful run cost, and where the underlying artifacts live.
The raw logs can stay raw. The receipt makes them findable and interpretable. Think about an ordinary expense receipt. It does not contain the restaurant's entire accounting system. It tells you who, when, what, how much, and enough detail to challenge the charge.
An agent receipt should do the same for delegated work. Who ran it? When did it run? What did it do? What did it change? What did it cost? Who accepted it? And where is the evidence if someone needs to challenge the result?
There are three common failure patterns to look for. The first is the activity receipt. It says the agent completed twelve steps, called five tools, and processed eighty files. That sounds busy. It does not say whether the outcome was useful.
An agent can execute every planned step and still solve the wrong problem with impressive discipline. Activity explains motion. It does not prove value. The second is the evidence orphan. The logs exist. The artifacts exist. The approval record exists.
But nobody owns the decision to accept or reject the result. So the evidence becomes a museum exhibit. Interesting. Detailed. And somehow not responsible for anything happening next. Every receipt needs an outcome owner and an exception owner.
The outcome owner says whether the work met the bar. The exception owner decides what happens when it did not. Those may be the same person. But the names have to exist before the incident. The third failure pattern is the immortal approval.
The workflow was approved once, so the organization treats that decision like a family heirloom. Meanwhile, the model changes. The tools change. The connected data changes. The policy surface changes. The license changes. And the approval keeps smiling from a document created when the agent still needed supervision to open a spreadsheet.
A receipt should record the policy version and the review date. It should also identify the events that force a new review. A model change. A new tool. A new data source. A new action permission. A material cost shift.
A failed quality threshold. Or a change in the control plane that affects logging, blocking, or inventory. Approval is not permanent context. It is a decision made under conditions. When the conditions move, review the decision. So here is the action for this week.
Pick one AI workflow that claims to complete work. Not your entire AI program. One workflow. Give yourself forty-five minutes. Build a one-page receipt for its most recent successful run. First: Name the accepted outcome. What did done mean?
Who decided the quality bar was met? Second: Record the execution path. Which model, identity, environment, data, tools, files, and delegated agents were involved? Third: Record the authority boundary. What could the agent read? What could it change?
Which actions required approval? Fourth: Record the economics. Usage, runtime, review time, retries, corrections, and rework. Fifth: Record the dependability result. Ready to use. Needs correction. Needs escalation. Sixth: Record the evidence. Logs, artifacts, source versions, approval notes, changed records, and the place another reviewer can reconstruct the run.
And seventh: Name the stop owner. Who can pause the workflow, revoke access, change the policy, or roll back the result? Then ask one uncomfortable question. Could a person who was not in the room understand what happened from this receipt alone?
If the answer is no, the workflow is not ready to scale. It may be ready to pilot. It may be ready to assist. It may be useful with direct supervision. But do not call it governed completion if the evidence disappears the moment the operator closes the tab.
That is the Work Agent Receipt Check. One accepted outcome. One execution record. One authority boundary. One full-cost view. One evidence path. One human owner with stop power. The agent can say done. The organization still has to prove it.
If the work matters, keep the receipt. That is it for this week's AI Change Desk. One episode. One emerging change. One useful action for the commute. I am Michael. Thank you for listening, and thank you again for the room to reset.
I will see you next week.