Full transcript
Delegation Quality Check
EP024 · May 6, 2026 · 14m 27s
Here is the uncomfortable thing about delegating work to an AI agent. The agent does not know which part of the task is normal Tuesday, and which part is please do not accidentally send this to a client.
It just sees the assignment. Make the deck. Clean the inbox. Build the model. Screen the K Y C file. Close the books. Very casual. Just letting software walk around with a clipboard and a tiny blazer. And if that sounds useful, it is.
Most of the time. If that also sounds terrifying, welcome. You are hearing it correctly. Because the new question is not: Can the agent do the task? The better question is: Who checks the handoff? Who checks the source?
And whether it is actually a source? Who checks the assumptions? And whether the numbers can survive a meeting? Who checks the permissions? Who checks the final artifact? And who says: this is ready to leave the room?
Or: no, it is not ready yet. That is the delegation quality problem. Not access. Not just trust. Quality. The part where useful work becomes risky because it looks finished. This week, the signals are loud. Microsoft is pushing Copilot Cowork further from conversation into action.
Agent 365 is becoming the control-plane story for agents. And Anthropic is packaging financial-services agents for work like pitchbooks, K Y C screening, valuation review, and month-end close. That is not three vendor announcements. That is one operating question wearing three different jackets.
When agents do the work, who certifies the work? Quick disclosure before we get into it. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led.
This is operational guidance, not legal advice. These are my opinions, and are not representative of any organization. I am source-checking this on May fifth, twenty twenty-six. Yes, the revenge of the fifth. We are allowing one nerd note, and then we are behaving like adults.
The timing matters. Some of this news landed today. Some of it is a few days old. And all of it points in the same direction. Agents are not just answering questions anymore. They are being handed work.
Welcome back to AI Change Desk. I am Michael. Today is the Wednesday episode for May sixth, twenty twenty-six. And this is episode twenty-four: Delegation Quality Check. The question today is simple. When AI does the assignment, who owns the review?
This connects directly to episode twenty-two and episode twenty-three. Episode twenty-two was the access-lifecycle check. The question there was: when the doors move, who updates the map? Episode twenty-three was the trust-boundary check. The question there was: when AI becomes infrastructure, who owns the boundary?
Today is the next layer. Because you can have the right door, and the right boundary, and still have a bad handoff. You can approve a tool, secure the account, log the agent, and still ship a deck with the wrong assumption sitting quietly on slide seven.
Slide seven is where mistakes go to start a family. That is why delegation quality matters. It is not the glamorous part of AI adoption. Nobody is putting a giant banner on a keynote slide that says: we have a pretty good review checklist.
But that is the difference between acceleration, and a closet full of operational mess. And yes, everyone has that closet. Useful is not the same as ready. Fast is not the same as approved. And generated is definitely not the same as governed.
First signal. Microsoft is making delegation feel more normal. On May fifth, Microsoft published a Copilot Cowork update. The frame is very clear. Cowork is about moving beyond chat and into execution. Microsoft says people are using it to orchestrate inbox workflows, conduct deep research, generate structured documents, and build full web pages.
It is built on Work I Q, which Microsoft describes as the intelligence layer that understands your data, your tools, and your organization. Then the update gets practical. Cowork is coming to iOS and Android through the Frontier program.
So delegation is not just something you do at the desk. It becomes something you can hand off from your phone, between meetings, on a commute, or while standing in the kitchen pretending the meeting is not still happening.
A bold era for people saying, I will circle back, and then assigning a robot to actually do the circling. Microsoft also talks about Cowork Skills. A skill is a reusable set of instructions for how a workflow should be done.
Structure. Tone. Process. The point is consistency. You capture the way the work should happen, then ask Cowork to apply it again. And Cowork plugins connect that work across documents, data, and line-of-business systems. Fabric I Q with Power B I.
Dynamics three sixty-five. Upcoming connectors like L S E G, Miro, monday dot com, and S and P Global Energy. Here is the operator translation. The interface is getting easier. The work is getting more connected. The agent is getting closer to the actual artifact.
That means the review cannot live only at the prompt. It has to live at the handoff. Do not just ask what the agent can do. Ask what the organization will accept as done. This is where Agent three sixty-five matters.
Microsoft says Agent three sixty-five is now generally available. At this point, I am mostly grateful it is not another thing called Copilot. Knock on wood. Branding has heard us before. The important idea is not the brand name.
The important idea is the control-plane shape. Observe. Govern. Secure. That is the language. And it tells you where the enterprise problem is going. Not just: which agents exist? But: where do they run? Which identities are attached?
Which data can they reach? Which tools can they call? What are they doing? And what happens when they do something weird at machine speed? Some of the Agent three sixty-five capabilities are generally available. Some are public preview.
Some are coming in June. That distinction matters. A control-plane announcement is not the same as a fully implemented local operating procedure. A dashboard is not a governance program. It is a dashboard. Beautiful rectangles. Still rectangles. But the direction is important.
Microsoft is explicitly talking about local agents, SaaS agents, cloud agents, network controls, registry sync, runtime alerts, and policy-based guardrails. That is not because agents are cute. It is because agents create exposure. And exposure needs ownership. If an agent can act, it needs an owner.
If it can access data, it needs a boundary. If it can produce work, it needs review. Second signal. Anthropic moved the same story into financial work. On May fifth, Anthropic announced agents for financial services. Ten ready-to-run templates.
Pitch builder. Meeting preparer. Earnings reviewer. Model builder. Market researcher. Valuation reviewer. General ledger reconciler. Month-end closer. Statement auditor. K Y C screener. That list is not a toy list. That is not: make me a birthday limerick about procurement.
Although, honestly, procurement probably deserves one. This is high-consequence work. Client materials. Models. Controls. Compliance files. Audit readiness. Books of record. Anthropic says each template packages three things: skills, connectors, and subagents. The skills define the task and domain knowledge.
The connectors provide governed access to data. The subagents handle specific subtasks, like comparables selection or methodology checks. That is useful. It is also a governance flare. Because the more complete the package gets, the more finished the output can look.
And the more finished the output looks, the easier it is for a tired human to assume the process was sound. That is the danger zone. Not because the agent is useless. Because it is useful enough to be trusted too quickly.
To Anthropic's credit, their announcement keeps humans in the loop. They say users review, iterate on, and approve Claude's work before it goes to a client, gets filed, or is acted on. That line matters. Because for operators, human in the loop cannot be a decorative phrase.
It has to mean a named role, a review artifact, a retained evidence path, and a stop condition. If nobody knows what the reviewer is reviewing, you do not have human oversight. You have a vibes-based turnstile. And the turnstile is wearing a fleece vest that says innovation.
So what do we do with this? We stop treating delegation as a productivity feature. We treat it as a quality system. Every delegated AI workflow needs four controls. First: the assignment boundary. What is the agent allowed to do?
What is it explicitly not allowed to do? Can it draft? Can it edit? Can it send? Can it file? Can it change a record? If the answer is unclear, the agent is not delegated. It is wandering.
Second: the source boundary. What data did it use? Was that data approved for the task? Was it current? Was it complete? Was it allowed to cross into this workflow? An agent that uses the wrong source politely is still using the wrong source.
Polite wrong is still wrong. Third: the review boundary. Who reviews the work? What are they checking? Accuracy? Compliance? Tone? Math? Source lineage? Client sensitivity? A generic approval checkbox is not enough. That is just a tiny ceremonial button.
Fourth: the release boundary. When can the artifact leave the conversation? When can it enter a deck, a memo, a ticket, a client email, a record system, or a decision process? That is the moment that matters. Because inside the chat, bad output is annoying.
Outside the chat, bad output becomes operational risk. The next-week action is simple. Pick one AI-assisted workflow. Just one. Not the whole enterprise. Not a three-hundred-line policy matrix that makes everyone suddenly interested in lunch. One workflow. By Wednesday, May thirteenth, twenty twenty-six, run a thirty-minute delegation-quality review.
Answer nine questions. What task is being delegated? Who is allowed to delegate it? Which source systems can the agent use? What artifact does the agent produce? Who reviews it? What evidence gets retained? What is the fallback if the agent fails?
What condition stops the workflow? And what plain-language message do users need before they try it? That is it. One workflow. Nine answers. Thirty minutes. If you cannot answer those questions for one agent workflow, you are not ready to scale it.
You are ready to create a very confident mess. The headline this week is not that agents are getting better. They are. The headline is that delegation is becoming normal. From Microsoft Cowork, to Agent three sixty-five, to Anthropic's finance templates, the same pattern is showing up.
Work is moving from human hands, to agent hands, and then back into human approval. That middle space is the danger zone. That is where assumptions hide. That is where sources drift. That is where a draft becomes a deliverable before anyone admits it happened.
So keep the frame clean. Access tells you who can enter. Trust boundaries tell you where the agent can operate. Delegation quality tells you whether the work can leave the room. Useful is not ready. Ready means reviewed.
Ready means evidenced. Ready means owned. That is the delegation-quality check. And that is where I would put the attention this week.