Somewhere this morning, somebody gave an A I agent a longer leash. Not a dramatic leash. Not the cinematic kind where the robot turns around, looks into camera, and says something ominous about humanity. I mean the boring enterprise kind. A longer task window. A new connector. A remote session. A stronger coding model. A permission that used to be temporary, but is now somehow called standard access. And the system probably looked fine. The dashboard was green. The demo was clean. The vendor post had screenshots. The release note had just enough words to make everyone feel like they understood what changed. Which is a dangerous amount of understanding. Just enough to approve something. Not enough to govern it. Because the moment an agent can work farther away from direct human supervision, you have a new question. Not: can it do the thing? That is the easy question. That is the confetti question. That is the question that gets answered in a product demo by someone with perfect Wi-Fi and suspiciously clean sample data. The better question is: what proof comes back with the work? If the leash gets longer, the receipt needs to get better. That is today's episode. Agent Reliability Evidence Check. Which sounds like a course you take after accidentally giving a raccoon admin access, but honestly, that raccoon has been quiet lately. Too quiet. Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions, and they are not representative of any organization. As I am source-checking this on Monday, June first, twenty twenty-six, the timing matters. OpenAI's latest official ChatGPT release-note movement I am using here is dated May twenty-ninth. Anthropic's Claude Opus four point eight announcement is dated May twenty-eighth. And AWS has fresh late-May implementation material on observability, agent test suites, and managed experiment workflows. So this is not a rumor episode. It is not a vibes episode. It is not me squinting at three screenshots and declaring a new era. We are going to keep the claims tied to dated sources. Then we are going to ask the operator question. If agents are getting more reach, more time, more autonomy-shaped workflow, and more ways to act across tools, what evidence should come back before anyone calls that reliable? Welcome back to AI Change Desk. I am Michael. Today is Monday, June first, twenty twenty-six. This is episode twenty-nine: Agent Reliability Evidence Check. The operating question is simple. When the agent can do more, what proof do you require before you trust the work? Let's start with the shape of the week. The last two episodes matter here. Episode twenty-six was about agent toolchain ownership. The question there was: who owns what the agent can reach? Because the agent is only as governed as the tools, connectors, permissions, and data pathways around it. Episode twenty-seven was about no-new-delta verification discipline. The question there was: if the official source did not move, why did your claim move? That episode was about not inventing novelty just because the calendar was hungry. Today connects those two. Because now we have a real source-backed agent-control story again. Not one giant story. Not a single press release that explains everything. More like a pile of small operating signals pointing in the same direction. OpenAI is moving Codex deeper into remote and delegated work patterns. Anthropic is describing a stronger agentic coding and workflow model with explicit control language around effort and behavior. AWS is publishing practical material around observing, testing, and managing L L M and agent workflows. Those are different vendors. Different layers. Different incentives. But for an operator, they rhyme. They all point at the same uncomfortable truth. A I work is becoming less like a one-shot answer, and more like a shift. A task run. A delegated process. A remote piece of labor with logs, state, permissions, cost, quality, and failure modes. The unit of governance is moving from the prompt to the work session. That is the turn. Not the model. Not the chat box. The work session. And work sessions need evidence. Let's talk about the OpenAI side first. The May twenty-ninth ChatGPT release-note entry includes several Codex-related movements. The details matter less than the pattern, but the pattern is hard to miss. There is Windows computer use in Codex. There is remote control over tasks. There are usage profiles. There are workflow improvements around sessions, terminal visibility, and the way users interact with work that is already running. There is also continuing connection tissue around repositories, M C P, and development workflows. Now, I am not going to turn that into a universal claim. I am not saying every organization has every capability today. I am not saying everyone should turn it on. I am not saying your dev team should put Codex in a tiny swivel chair and call it an intern. Please do not do that. The chair budget alone would be unacceptable. The operating point is simpler. The interface is moving from asking a model something, to managing work a model is doing. That is a different control problem. When a model answers a question, you review the answer. When an agent works a task, you review the task path. What did it touch? What did it change? What did it skip? What did it assume? What did it spend? Where did it get stuck? And who was supposed to notice? That last question is the one that quietly ruins everyone's afternoon. Who was supposed to notice? Because a lot of A I governance still behaves like the main risk is the final answer. The output. The summary. The code diff. The slide. The support draft. The thing someone can look at and say: hmm, that feels wrong. But agent workflows create middle risk. The risk happens in the path. The selection. The tool call. The retry. The hidden assumption. The permission that worked once, so everyone stopped asking why it worked. If your review only sees the final answer, your control surface is already too small. This is where the Anthropic signal fits. Anthropic's Claude Opus four point eight announcement is also a capability story, but I want to read it as an operating story. Anthropic describes stronger performance on software engineering benchmarks. They describe better behavior in longer, more complex work. They talk about effort control, dynamic workflow behavior, system entries in the messages array, and an honesty improvement metric. Those are vendor-reported claims, so keep the attribution clean. But the categories are useful. Performance. Effort. Workflow behavior. Instruction structure. Honesty. That is not just a model card. That is an operating checklist trying to escape the product announcement. Because once a model is strong enough to do more work, the question becomes: how much work should it be allowed to attempt before the organization asks for a checkpoint? How much effort is appropriate for the task? How much cost is appropriate for the task? How much uncertainty is acceptable before the agent should stop, ask, or escalate? And this is where a lot of teams get cute. They say, we trust the model. Which is a sentence that should be placed gently on a table, covered with a towel, and not allowed near production systems. You do not trust the model. You trust a workflow under conditions. You trust it on this data, with this tool access, for this task class, under this review gate, with this rollback path, and this human owner. Trust is not a personality trait. It is a control state. That sounds sterile. I know. It sounds like something printed on a laminated card in a conference room named after a tree. But it is also the difference between useful A I adoption and a very expensive magic trick. The magic trick says: look, it worked. The operating system says: show me why it worked, where it might fail, and what happens when it does. Now let's bring in AWS, because this is where the abstract part gets less abstract. AWS has late-May implementation material around L L M observability for SageMaker A I inference, including infrastructure health and output quality. They also have a piece on evaluating deep agents with LangSmith on AWS. And another on building a test suite that grows with an agent through dataset management in Bedrock AgentCore. That is a mouthful. It sounds like someone spilled a cloud architecture diagram into a bowl of alphabet soup. But underneath the naming pile, the operator point is useful. Agent reliability is not a feeling. It is an evidence stack. Traces. Metrics. Logs. Regression tests. Experiment records. Model and prompt versioning. Workflow comparison. Fallback decisions. This is the part teams skip because it is less fun than the demo. Nobody claps for the trace ID. Nobody posts a launch thread that says: we are thrilled to announce our retry policy now has adult supervision. Although honestly, I would respect that thread. I might even like it. But this is what separates a toy workflow from an operating workflow. A toy workflow can impress you once. An operating workflow can explain itself after the third weird Tuesday. And there is always a third weird Tuesday. Reliability is not when the demo works. Reliability is when the failure leaves enough evidence to fix the system. So what does that mean for an organization this week? I would not start with a giant governance program. Please do not respond to every agent announcement by creating a seventeen-tab spreadsheet called enterprise A I operating model final final revised. That way lies madness. And probably conditional formatting. Start smaller. Start with five receipts. Every agent workflow should produce five receipts before you call it reliable enough to scale. First: the scope receipt. What can the agent reach? Not in theory. In practice. Which repository? Which folder? Which browser session? Which ticket queue? Which customer record? Which calendar? Which third-party tool? Which M C P server? Which internal system? If the answer is, well, it depends, then congratulations, you have found the first control gap. It does depend. That is why you write it down. Scope is not a vibe. Scope is an inventory. And if the agent can reach a thing, that thing needs an owner. No owner, no reach. That is the rule. Second: the effort receipt. How long can the agent work? How much compute can it spend? How many retries are allowed? How many files can it inspect? How many actions can it attempt before a human checkpoint? This matters because effort is where cost and risk hide together. A cheap-looking task can become expensive if the agent keeps trying. A low-risk task can become high-risk if the agent keeps broadening the search. A useful assistant can become a small caffeinated raccoon if nobody knows when it is supposed to stop. And yes, apparently the raccoon is back. It has discovered retries. Set effort limits by task class. Not by general enthusiasm. A documentation cleanup can have one threshold. A code migration can have another. A customer-facing communication should have a tighter human checkpoint. A workflow touching regulated data should have a very different path. Effort is a control. Treat it like one. Third: the quality receipt. What proves the output is good enough? For code, that might be tests, lint, review, security checks, and a rollback plan. For content, that might be source traceability, claim review, legal language, brand tone, and human approval. For operations, that might be a runbook comparison, a second-system check, or a required sign-off before execution. The key is not to say: we will review it. That is not a quality system. That is a hope wearing a badge. Say what review means. Who reviews? Against what standard? With what evidence? Before what action? And what happens if it fails? This is where test suites for agents matter. Not because every organization needs the same toolchain. Not because one vendor blog solves your process. But because the concept is right. If an agent will run a workflow repeatedly, you need regression tests for the workflow. You need examples of known-good behavior. Known-bad behavior. Edge cases. Permission failures. Data quality failures. Escalation moments. The boring stuff. The useful stuff. If you cannot test the workflow, you are not ready to scale the workflow. Fourth: the drift receipt. What changed since the last good run? This is the sneaky one. Because an agent workflow can be reliable on Monday, weird on Wednesday, and quietly dangerous by Friday, without anyone changing the headline feature. A dependency changed. A prompt changed. A model version changed. A model retirement date appeared. An editing surface moved. A data source changed. A permission changed. A schema changed. A policy changed. A user changed the instructions because the old ones were annoying. Which is the most human root cause imaginable. And this is not hypothetical. The same OpenAI release-note page also includes a May twenty-eighth model-lifecycle signal. G P T four point five has a June twenty-seventh ChatGPT sunset. O three has an August twenty-sixth ChatGPT sunset. And canvas is no longer part of the newer G P T five point five Instant and Thinking path. That is not the lead story today. But it belongs in the fallback receipt. If a workflow depends on a model, a mode, or an editing surface, the retirement date is not trivia. It is operational evidence. This is where observability matters. Logs are not just for debugging. Traces are not just for engineers. Metrics are not just for the dashboard people, who already have enough laminated anxiety in their lives. Evidence is how you know whether the work session is still the same kind of work session. For agent workflows, you need a simple drift view. Last good run. Current run. Changed inputs. Changed tools. Changed model or prompt. Changed result. Changed cost. Changed human intervention. Drift is not only model behavior. Drift is the whole operating context moving while everyone pretends the workflow stayed still. Fifth: the fallback receipt. Who stops it? Who reroutes it? Who explains it? Who tells users what to do when the agent is unavailable, wrong, slow, expensive, or no longer approved? This is the receipt nobody wants to write, because fallback planning makes the shiny thing look less magical. But magic is not a business continuity plan. If your fallback plan is, use another model, that is not a plan. That is a bumper sticker with a login screen. A real fallback plan names the task, the owner, the alternate path, the user message, the data handoff, and the point where the team stops trying to make the agent work and uses the old process. There is no shame in the old process. The old process may be clunky. It may involve spreadsheets. It may involve a human named Denise who knows where the weird form lives. But if Denise is the fallback, Denise deserves to know before the system breaks. A fallback that depends on a person who was never told is not a fallback. It is a surprise meeting. Now, there is a management trap here. The trap is turning these receipts into theater. A nice template. A dashboard. A status field. Five green checkmarks that make everyone feel calmer while the actual workflow remains weird. So do not build the receipt system as paperwork first. Build it around decisions. The scope receipt decides whether access is allowed. The effort receipt decides when the task must pause. The quality receipt decides whether output can ship. The drift receipt decides whether the workflow still belongs in the approved lane. The fallback receipt decides what happens when the answer is no. That is the difference. Evidence that does not change a decision is decoration. Evidence that changes a decision is governance. Do not collect receipts for the scrapbook. Collect receipts for the stop sign. There is also a leadership translation problem here. Because if you walk into a leadership meeting and say, we need better agent observability, somebody will nod politely while mentally checking whether the next meeting has lunch. That is not because they are careless. It is because observability sounds like plumbing. And plumbing only becomes interesting when the ceiling is wet. So translate the evidence conversation into management language. Not: we need traces. Say: we need to know what the system did before we approve more reach. Not: we need regression tests. Say: we need to know whether the workflow still passes the examples that matter to the business. Not: we need prompt versioning. Say: we need to know whether the instructions changed before the results changed. Not: we need rollback. Say: we need a named path back to the human process when the agent is wrong, slow, or unavailable. This is not wordsmithing. It is how you get the right people into the decision. Security hears reach. Finance hears effort and cost. Legal hears claims and records. Operations hears fallback. Product hears user impact. Engineering hears tests, logs, and failure reproduction. Leadership hears whether the organization can scale the workflow without making tomorrow's incident report more creative than today's strategy deck. That is the audience map. If your agent review only speaks engineering, you will miss the operating risk. If it only speaks policy, you will miss the failure mechanics. If it only speaks finance, you will miss the quality problem until customers find it for you. And if it only speaks innovation, you are probably already using the word transformation too much. Please drink water. The evidence check gives every group a clean question. Security asks: what can it reach? Finance asks: what can it spend? Engineering asks: what proves it works? Operations asks: what changed? Support asks: what do we do when it breaks? Legal asks: what claim can we defend? Leadership asks: can this scale without creating an invisible dependency? That is why agent reliability is a management discipline, not just an engineering feature. The technology can get better and still make the organization worse if the operating model does not catch up. That is the part people underestimate. A stronger model can produce more work. More work can create more review load. More review load can create rubber-stamp approvals. Rubber-stamp approvals can create hidden drift. Hidden drift can create bad customer outcomes, bad compliance posture, or just a team that spends every Friday asking why the thing that worked on Tuesday is now acting like it was raised by raccoons. Again with the raccoons. I am trying to grow. They keep returning to the evidence. So the management move is not to slow everything down. It is to make permission conditional. You can use this workflow while the receipts are current. You can expand this pilot after the quality receipt passes. You can connect this tool after the scope owner signs off. You can run longer tasks after effort limits exist. You can publish output after the review gate is real. You can keep the agent in production while drift stays inside tolerance. Permission should expire when evidence expires. That is the sentence. If you remember nothing else from this episode, remember that. Permission should expire when evidence expires. So here is the action block for this week. Pick one agent workflow. Just one. Not every tool. Not the entire A I estate. One workflow where the leash has gotten longer. Maybe code review. Maybe research. Maybe support drafting. Maybe sales prep. Maybe internal reporting. Maybe the weird workflow everyone knows is happening but nobody wants to describe because then someone will ask who approved it. Start there. Run an agent reliability evidence check. Five questions. First: what can it reach? List the systems, data, tools, and owners. If you cannot list them, you do not approve expansion. Second: how much effort is allowed? Set task limits. Time. Cost. Retries. Files. Tool calls. Escalation points. If no limit exists, the limit is currently your credit card and your luck. That is a bad control pair. Third: what proves quality? Name the tests, review gates, source checks, or acceptance criteria. If the proof is only, someone will look at it, make that sentence illegal in the meeting. Not legally illegal. Meeting illegal. Very different court. Stricter snacks. Fourth: what drift do we watch? Model version. Prompt version. Tool access. Data source. Output quality. Cost. Human intervention. Complaints. Retries. Decide what movement would force a review. Fifth: what is the fallback? Name the human owner. Name the alternate workflow. Name the message to users. Name the point where you stop retrying and switch paths. Then write one sentence at the top. This workflow is approved only while these five receipts stay current. That is it. That is the start. Not a giant program. Not a new bureaucracy. A reliability check that attaches evidence to permission. The deeper point is this. A I agents are going to keep getting better at doing work away from the immediate attention of the person who asked. That is useful. That is also the risk. Not because the machines are plotting. Most of the time they are not plotting. Most of the time they are just confidently misunderstanding a badly framed task, which, to be fair, is also a thing humans do in meetings. The risk is that organizations confuse capability with reliability. Capability says: it can do more. Reliability says: we know when to trust it, when to stop it, and how to prove the difference. Those are not the same sentence. They should not go through the same approval door. This is why I keep coming back to operating questions instead of hype questions. The hype question is: which model is best? The operating question is: best for which workflow, under which limits, with which evidence, owned by whom, and reversible how? That is less exciting. It is also the part that keeps the building from filling with invisible dependencies and tiny unapproved robots holding clipboards. Which is a terrible mental image. But very accurate. So before you expand an agent pilot this week, ask for the receipts. Scope. Effort. Quality. Drift. Fallback. If those receipts exist, you have something to manage. If they do not, you do not have agent reliability. You have a really impressive demo with a payment method attached. Do not scale the demo. Scale the evidence. That is the work. Welcome to the less glamorous, more useful part of A I adoption. The part where the system has to explain itself after the applause. That is it for today's AI Change Desk. I am Michael. If this episode helped, send it to one person who is excited about agents, and one person who has to clean up after agents. They should probably meet. Before the raccoon gets another connector. Thanks for listening.