AI Change Desk | EP047: Who Checks the AI Safety Claim? Published September 28, 2026 | Runtime 19:02 Transcript of the published audio, including the reusable disclosure and opening/closing song voice lines. RSS.com transcription checked against the approved script; minor recognition and spelling errors corrected. Imagine it is Thursday afternoon. Your team plans to launch an AI assistant on Monday. An outside reviewer sends one sentence. We recommend delaying the launch until this test is complete. The project sponsor reads it and asks, is that a recommendation or do we have to stop? Nobody in the meeting can answer. The reviewer has expertise. The sponsor has the budget. The operations team has the launch button. Somewhere between those three facts, everyone assumed a decision had been assigned. The meeting invitation says governance. That is as far as the paperwork got. Now the person running the launch has a very practical problem. Whose instruction changes Monday? Before you tell a customer that independent experts reviewed your AI, establish what those experts could see, what they could say, and what your organization did with their advice. Because a recommendation only helps if someone owns the response. Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization. One additional disclosure as the desk evolves. My current role includes privacy work. so you will hear me pay closer attention to purpose, access, retention, deletion, and accountability. I will not discuss non -public work here. These are my personal views, and they do not represent the state of Oregon or any other organization. This is AI Change Desk. Thank you for choosing to tune in. Welcome to AI Change Desk. I am Michael. In episode 46, we asked who owns the incident after a system crosses a boundary. This time, move one decision earlier. Who checks the safety claim and who answers the finding before the organization commits? I am checking the public record on Monday, September 28th, 2026. The dates matter here. A new disclosure can describe older behavior, and an investigation can still be open after a safeguard changes. On September 22nd, OpenAI proposed principles for third -party safety assessments. The proposal calls for defined claims, suitable access, clear methods, and arrangements for reporting and responding to findings. It describes much of this work as longer -term, separate from launch schedules. That is, a proposal for scrutiny. It is not a completed assessment or an automatic release veto. Now, what happens after a troubling finding? On September 25th, OpenAI reported 53 instances in which user -provided images reached image hosting sites through its research agents. The company says the links were not publicly listed, and most of the content had been removed, with removal continuing. It says the cases preceded safeguards described in its technical report. It also says enterprise, business, and API data are excluded from training unless an administrator enables it. The wider investigation remains open. That is, the company's account, not an independently completed finding. 53 instances does not tell us the number of affected people. The disclosure date does not mean the activity happened that day. Here is the operating distinction I would carry into the room. Permission for one data use is not permission for every later destination. An unlisted link is not a privacy conclusion. And removal underway is not the same as a verified end state. Do not make the person whose information moved disappear behind the diagram. Their practical question is simpler than our architecture slide. Where did it go? Who could reach it? And what happens now? I am not presenting a new incident response checklist. We covered response authority in episode 46. This is about the sentence an organization can honestly say afterward. We changed a control. We tested that control. We verified a particular outcome. Those are three different claims. The evidence for one does not automatically establish the next. There is a practical follow -up today from the separate Australian incident we discussed in episode 46. Minister Katy Gallagher says Services Australia has changed its notification route. That reporting inbox now goes directly to a cyber center staffed around the clock. She also says investigations are continuing. That is a reported process change, not proof of a resolved incident. But it gives a reviewer something specific to test. Does the next serious warning reach someone who can act? That is a concrete change for an assessor to examine. Now the second supporting signal, a worked example from Microsoft. Its September 24th article describes testing a billing support agent, adding a runtime control, and testing again. The tool is called RunAssertEval. In one part of its cross -customer test, Microsoft reports violations fell from about 21 % to about 9%. In another, they fell from about 44 % to zero. Those are different parts of a test, not interchangeable scores. These are sample observations, not a promise about your deployment. There were still failures. Microsoft describes turning the method into a repeatable release gate as work ahead, not a completed capability. I have read the article. I have not independently run its code or reproduced its results. A team can make a real improvement and still have work it must not approve. Progress deserves credit. It does not need a certificate invented by the slide deck. Which finding changed the work? Which failure is still open? And who is allowed to accept that remaining uncertainty? A capability test asks what a system can do under specified conditions. A safeguard test asks whether a control limits a particular behavior. A deployment decision asks whether the proposed use is acceptable here. You can need all three. You cannot replace one with the logo attached to another. Think about a familiar claim. This assistant will only use approved information. Before accepting that sentence, you need to know which assistant, which information, which permissions, and which attempted violations were tested. Then you need someone to answer for the gaps. The goal is not to turn every buyer into a frontier model researcher. It is to stop a narrow finding becoming an unlimited reassurance as it travels from a technical report into a slide deck. By the time the slide says safe, the sentence explaining safe under which conditions may have disappeared. Let us keep that sentence in the room. Here is the point for the person who has to run the system. An advisor can help you notice something important. An evaluator can examine a defined claim and produce evidence. An authorized decision maker can accept, narrow, delay, or reject the proposed action. Those are working definitions for this episode. Your actual agreements may use different terms. Write down the job in plain language before relying on the title. Start with a sentence you could put in a launch meeting. This reviewer can recommend a delay, and this named person must answer before deployment. Notice what that sentence gives you. It does not pretend that every expert has a veto. It creates an obligation inside your own process to deal with expert advice. You could choose a stronger rule for a particular use. A specified unresolved finding could prevent deployment until an authorized reviewer clears it. Or you could allow a documented exception within a narrowly defined limit. Those are design choices to make deliberately. An organization should not discover its choice when somebody is already counting down to launch. The decision also needs an object. Approval to continue a small internal test does not answer whether the same system may serve customers. Advice on one model version does not automatically cover a replacement version with different tools. Give the reviewer and decision maker the same question. Name the system, the intended use, the population affected, and the action being considered. Otherwise, one person can be right about a narrow test, while another person repeats that answer as broad reassurance. And make room for disagreement. A useful record can say that a reviewer recommended waiting and that the decision maker chose a smaller pilot for a stated reason. That record should preserve the concern, identify the remaining uncertainty, and explain why the narrower action is acceptable under the organization's rules. It should also say what would cause the decision to change. If the person signing cannot explain that in ordinary language, adding the reviewer's name to the presentation will not fix the decision. There is value in expert advice, even when the expert cannot make the final call. The value depends on whether the organization has a clear way to hear it, answer it, and act on it. Let us stay with the Thursday launch meeting. This is a hypothetical example using invented facts and synthetic records. A customer support team wants an assistant to draft replies about account changes. For the initial pilot, employees would read and approve every reply before sending it. The assistant would not change an account directly. An outside reviewer has tested ordinary requests and found that the drafts usually stay within the supplied policy. But the reviewer has not tested requests where old policy documents conflict with current instructions. The recommendation is to delay the pilot until that conflict test is complete. The sponsor hears that the ordinary tests look good. The operations lead hears that a required test is missing. The reviewer believes they have advised against Monday's launch. All three people can leave the meeting with different memories unless someone records the decision. First, make the missing evidence concrete. Which conflicting documents were not tested? Which version of the assistant was reviewed? Could employees recognize the failure if it appeared in a draft? What evidence would answer that question? There may be a defensible way to narrow the pilot. Perhaps the team can remove the disputed document set use only synthetic accounts, and keep every output inside a training exercise. That would be a different proposal for the reviewer and decision maker to consider. There may also be no sensible narrower use. If the assistant's central task requires the disputed documents, removing them could make the pilot meaningless. In that case, postponing may be the useful decision. The episode cannot choose for that imaginary team. The point is to expose the choice rather than let the word reviewed conceal it. Now imagine a second problem. The reviewer was shown a prepared set of outputs, but could not inspect the documents the assistant received. That does not make the review worthless. It limits the claim it can support. The reviewer may have assessed the quality of those outputs. They may not have assessed whether access controls consistently limited the underlying information. Write that distinction beside the finding, where the launch owner will see it. Do not hide the access limitation in an attachment that never reaches the decision meeting. And if broader access is needed, arrange a controlled way to provide relevant evidence. Outside expertise is not a reason to hand over customer records, live credentials, or every log the organization holds. Use synthetic examples where they answer the question. Where they do not, have the appropriate owners decide what evidence can be shared, with whom, under what controls, and for how long. Now we can describe two possible outcomes clearly. One. The owner delays the pilot, names the missing test, and assigns a date to review the result. 2. The owner authorizes a smaller synthetic exercise, records the unresolved finding, and prohibits customer use until the remaining question is answered. Neither outcome says the reviewer approved something they did not approve. Neither relies on silence as consent. The practical failure would be a third outcome. Monday arrives, the original pilot launches, and the evidence folder contains a recommendation nobody answered. A review should leave the team with a more precise decision than it had before the review began. There is another distinction worth slowing down for. Independence has more than one dimension. For your own review arrangement, ask who pays, who chooses the question, who can obtain evidence, who can challenge a conclusion, and who controls communication of the result. Treat those as questions to investigate, rather than a formula that produces a certificate of independence. Payment can create a conflict to manage. No payment does not magically create access, decision rights, or a complete picture. Working inside a company may provide valuable context while still requiring clear boundaries around what can be reported. You do not have to resolve the entire debate about independent AI evaluation to improve one procurement or launch decision. Ask what the arrangement permits, what it excludes, and how those limits affect the claim you intend to make. A team might need a specialist to critique a research method. Another team might need a reproducible test of one failure mode. Another might need an authorized approval before a particular use can proceed. Commission the job you need, then describe the completed work accurately. If someone reviewed a draft plan, say that. If someone tested a specified version against specified cases, say that. If an authorized owner accepted a remaining risk, identify that as the owner's decision. Avoid letting the name of a respected institution do work that the evidence has not done, and make the response proportionate. Not every editorial suggestion needs an executive decision. A concern that changes the permitted audience, data access, or consequences of failure needs a clearer route than a comment about wording. Agree on those routes before the review begins. Let the reviewer know how to flag a consequential concern. and let the operator know what to do when one arrives. The aim is a process in which useful criticism can alter a decision, and a person taking a different course has to explain the choice. Here is the workweek action. Set aside 45 minutes with one proposed use, one reviewer, and the person who can actually make the decision. If the reviewer cannot attend, use their written finding and have a named person check your interpretation with them afterward. Do not present a role play as the reviewer's real agreement. Use a simple recommendation to decision check. Five questions. What is being decided? What was examined? What was recommended? Who must respond? And what changes next? This is an editorial exercise to test your process. It is not an external standard or a finding that your organization has met its obligations. For the first seven minutes, write the decision in one sentence. For example, may this specific assistant draft replies for a small internal training group using synthetic account records next week? Name the version, the use, and the limit. If the group cannot agree on that sentence, stay there. There is little value in reviewing three different proposals under one title. From minute 7 to minute 15, map the evidence the reviewer actually had. Write down what they examined and what they could not examine. Separate the test provider from the person who actually ran it. Note whether the controls match the intended deployment. Include the important conditions without copying sensitive source material. Ask whether a missing item changes the decision. A reviewer does not need to see everything. They need enough relevant evidence for the particular claim you plan to rely on. Mark an unknown as unknown. Give it an owner and a date. Do not turn an empty field into an assumed pass because the meeting is moving quickly. From minute 15 to minute 23, introduce a disagreement. Use the reviewer's actual unresolved concern if there is one. Otherwise, create a clearly hypothetical concern that would matter to the proposed use. For our support assistant, it is the missing test of conflicting policies. Have one person state the recommendation precisely. Then, have the decision maker explain what their role allows them to do with it. Can they postpone? Can they narrow the pilot? Does another authority have to clear this class of concern? Is an exception permitted, and who may approve it? The important discovery may be that the person everyone expected to decide does not hold that authority. Better to learn that during the exercise. From minute 23 to minute 33, write the response. Choose a disposition in ordinary words. Accept the recommendation, choose a narrower action, ask for more evidence, or decline the recommendation with a reason that the proper authority can stand behind. Your organization's rules may rule out some of those options. The exercise does not create an exception power that nobody has. Record the reviewer, the responding owner, the reason, and any unresolved condition. Preserve the reviewer's original concern rather than rewriting it to sound like agreement. If the reviewer disputes your summary, keep that disagreement visible and correct any factual error. A useful process does not require everybody to sound pleased. From minute 33 to minute 40, trace the response into the actual work. If the decision is delay, who changes the launch plan? If it is a narrower pilot, who limits the audience and data? If another test is required, who runs it and who checks the result? Look for the place where the decision becomes an instruction that someone can execute. A signed document can still leave Monday's launch unchanged if nobody updates the work. Use a safe tabletop or a controlled test environment. Do not change production permissions just to make the exercise feel realistic. For the last five minutes, set the reopening condition. What change would require a fresh decision? A new model version, a new source of data, a broader audience, or an unresolved finding that survives its deadline could each matter. Choose the conditions relevant to your use. Name who notices the change and who brings the decision back for review. At the end, read the five answers aloud. The decision, the evidence, the recommendation, the responding authority, the change to the work. If one answer is missing, your next action is to resolve that gap with a named owner. 45 minutes should expose the gap. It cannot manufacture evidence or authority you do not have. By Friday, the person operating the assistant should be able to explain why the permitted use is different or why it stayed the same after the review. They should know which concern remains open, who answered it, and what would bring the decision back. That is a more useful result than a presentation that says experts were consulted. The new assessment proposals and test disclosures give us a reason to examine this now. The practical test belongs inside our own organizations. Can an outside recommendation reach an accountable decision? And can that decision reach the work? For the next launch, procurement, or material change, take one recommendation and follow it all the way through. Keep the expert's finding intact. Keep the decision maker visible. Keep the resulting action specific. Expert advice deserves an answer. The person making the decision still needs to own it. I am Michael. This is AI Change Desk. This is AI Change Desk. See you next time.