Full transcript
Model Routing Check
EP021 · Apr 27, 2026 · 27m 23s
Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.
As I am source-checking this on April twenty-sixth, twenty twenty-six, the timing matters. The model news is fresh. The infrastructure news is fresh. And the May eighth remediation date is close enough that it should already be on somebody's board.
Intro music
Here is the uncomfortable thing about an AI upgrade: It does not always arrive like a software release. Sometimes it arrives like a new traffic pattern in a city that forgot to update the signs. At nine in the morning, one team has the better model.
At ten-thirty, another team is still on the old path. By lunch, the sales deck sounds sharper, the legal prompt gets longer, the coding assistant starts taking scenic routes through your repo, and finance is staring at usage like the meter just learned to freestyle.
Then somebody asks: "So, are we live on the upgrade?" That sounds like a yes-or-no question. It is not. It is five questions wearing one very confident jacket. Who gets access? What is the fallback? Which prompts were retested?
Who owns the patch deadline? And where is the evidence when this gets reviewed? That is why this episode matters. If you lead operations, security, legal, finance, product, or change work, the risk is not that a stronger model exists.
The risk is that a stronger model quietly changes the route of the work before the organization changes the controls. That is the operating problem this week. Not whether the model is impressive. It is. Not whether the vendors are moving fast.
They are. The question is simpler and more uncomfortable: Who approved the route the work now takes? Welcome back to AI Change Desk. I am Michael. Today is the Monday main episode for April twenty-seventh, twenty twenty-six. And this episode connects directly to the last two.
Episode 19 was about release gates. Stronger models, cyber boundaries, task budgets, capacity, rollback, all the grown-up machinery that makes an upgrade survivable. Episode 20 took that same gate and pointed it at visual artifacts. Because a picture is not just a picture once it becomes public, branded, reused, exported, or handed to another team.
Episode 21 is the routing check. If episode 19 asked, "Do you have a gate before the smarter thing scales?" And episode 20 asked, "Do you have a manifest before the prettier thing ships?" This one asks: When the model path changes mid-week, do your controls still know where the work is going?
That is the thread. Not model hype. Not benchmark karaoke. Not a spiritual retreat for product announcements. A routing check. Because a stronger model does not only change output quality. It changes who gets access, what fallback means, which prompts still behave, how review load moves, how spend appears, how security deadlines land, and whether your standards evidence is real or just a file named "final final v three."
Which is, historically, not a file you should trust with civilization. We have three signals. First: OpenAI published GPT five point five on April twenty-third, and updated the post on April twenty-fourth to say GPT five point five and GPT five point five Pro are available in the API.
Second: Anthropic and Amazon announced an expanded compute collaboration on April twentieth, securing up to five gigawatts of capacity for training and deploying Claude. Third: OpenAI still has a hard May eighth remediation date from the Axios developer tool compromise response, and NIST's AI R M F critical infrastructure concept note gives operators a clean way to turn standards talk into owner, evidence, and due-date work.
Three stories. One management question. If access changes, capacity concentrates, deadlines stay real, and evidence is still fuzzy, who owns the route? Story one. GPT five point five is a model-routing problem. OpenAI's April twenty-third post introduces GPT five point five as a stronger model for real work.
OpenAI describes it as useful for writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. That sentence should make operators sit up. Because "moving across tools until a task is finished" is not a cute product phrase.
That is a workflow boundary walking around with shoes on. That walks around like it owns the place, and probably your coffee pot. The model is not only answering. It is planning, using tools, checking work, navigating ambiguity, and continuing.
That can be extremely useful. It can also turn a casual upgrade into a production routing change before anyone has finished their coffee or your coffee while it is at it. And yes, the model announcement includes the normal performance language.
Better coding. Better knowledge work. Better tool use. Better long-running work. Fewer tokens on some Codex tasks. Strong safeguards. Rollout to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex. GPT five point five Pro to Pro, Business, and Enterprise in ChatGPT.
API availability added in the April twenty-fourth update. All of that matters. But the operator read is not "new model good." The operator read is: Does your organization know when this model should be used, when it should not be used, what happens when it is unavailable, and who owns the fallback?
Because the first week of a stronger model is where teams accidentally invent shadow routing. One group gets access and starts using it. Another group cannot access it yet and quietly falls back. A third group changes the model because "the new one is better."
Someone runs a critical prompt and gets a materially different answer. Someone else gets a cleaner answer with fewer citations. Another workflow suddenly costs more, or less, but nobody knows why because the dashboard is a painting of a dashboard.
Very modern. Very laminated. Completely unhelpful. This is where I want to slow down. The dangerous assumption is that a model upgrade is a single event. It is not. It is a routing event. Who has access? Which team is eligible?
Which plan or tier applies? Which model is primary? Which model is fallback? Which workflows require citation? Which workflows require human approval? Which prompts need regression testing? Which outputs are allowed to be more autonomous? Which outputs are not?
Who decides when the route changes? That is the real work. And it is not glamorous, which is why it gets skipped. Nobody wants to be the person in the meeting saying, "Before we celebrate the new model, can we talk about fallback behavior?"
That person is not invited to many parties. But that person saves the organization from discovering production drift through customer complaints, confused reviewers, surprise usage, and one extremely tense Friday calendar invite. Here is the ordinary failure mode.
A team has ten critical prompts. Legal summary. Customer reply draft. Sales proposal. Engineering review. Security triage. Finance analysis. Research brief. Support escalation. Executive summary. Public copy. Those prompts worked well enough last week. Then the model path changes.
Nobody reruns the top prompts across the old and new paths. Nobody checks whether the citation behavior changed. Nobody measures correction load. Nobody asks whether the fallback answer is now worse, slower, more verbose, or just more confident while being wrong in a new font.
And now the team is arguing about taste. "This sounds better." "This sounds flatter." "This sounds more complete." "This sounds too cautious." "This sounds expensive." Those are not just taste claims. Those are measurement gaps wearing little mustaches.
The question should be: What changed in the route? Did the primary model change? Did the fallback change? Did access change? Did effort level change? Did the tool path change? Did the review standard change? Did the prompt harness still pass?
Did correction load go up? Did approval time go down for the right reasons, or because reviewers stopped reading? That last one is not a joke. Sometimes better output creates less review. Sometimes less review is earned. Sometimes less review is just everyone being seduced by formatting.
Formatting is not governance. Formatting is a nice outfit. So the first operating rule is: Do not approve the model. Approve the route. Approve the model for a workflow, with a fallback, with a reviewer, with a measurable failure condition.
If the new model is used for engineering work, name the test requirement. If it is used for customer-facing output, name the approval requirement. If it is used for legal or compliance drafting, name the citation requirement. If it is used for research, name the source standard.
If it is used for agentic computer work, name the stop authority. That is the difference between adoption and drift. Adoption has a route. Drift has vibes and screenshots. Story two. Compute concentration is now a continuity question.
On April twentieth, Anthropic announced an expanded collaboration with Amazon. Anthropic says it signed a new agreement to secure up to five gigawatts of capacity for training and deploying Claude. It also says more than one hundred thousand customers now run Claude on Amazon Bedrock.
The announcement includes new Trainium capacity coming online, a longer-term commitment to AWS technologies, and Claude Platform availability through AWS coming soon. Anthropic also says Claude remains available to customers on AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure Foundry.
This is where the normal business-news reading is too small. The story is not only: large vendor signs large infrastructure deal, everyone nods like they understand electricity. The operator read is: If your workflow depends on one model family, one provider path, one region, one procurement route, one identity setup, or one cloud integration, do you understand the continuity assumption you just made?
Because stronger models do not float in the air. They run somewhere. They depend on capacity, chips, regions, contracts, enterprise controls, billing paths, identity policies, and vendor roadmap decisions. And once a team builds a workflow around that path, "we can always switch later" becomes one of the great bedtime stories of technology management.
I love a comforting story. I do not love one that wakes up as a migration project. This is where episode 16 comes back. Episode 16 was about capacity as an operating constraint. Not abstract national capacity. Real planning capacity.
Can we get access? Can we keep access? Can we afford access? Can we route around failure? Can we degrade gracefully? Can we explain to the business which workflows depend on which provider path? The Anthropic-Amazon signal makes that less theoretical.
If Claude-heavy operations deepen inside AWS, that may be good for governance, billing, identity, and procurement alignment. It may also create practical dependence. Both can be true. That is the part operators need to hold without getting dramatic.
The question is not: Is this vendor good or bad? The question is: What are we depending on, and did we write it down? For a lot of teams, the answer is: we are depending on it emotionally.
Which is not a continuity plan. It is a mood board with an invoice. Here is the practical continuity check. Pick the AI workflow that would hurt the most if its model path changed this week. Now answer six questions.
What model family does it use? What cloud or platform path serves it? What identity and billing controls wrap it? What fallback exists if access degrades? What output quality is acceptable under fallback? Who is allowed to decide that fallback is now in effect?
If nobody can answer those questions, the problem is not the vendor. The problem is that your workflow has become dependent before your operations record caught up. And that is the quiet pattern across this whole episode. The technology moves first.
The workflow follows. The documentation jogs behind it in dress shoes. By the time leadership asks for the map, the map is an oral tradition. Oral tradition is beautiful in culture. It is less beautiful in model operations.
So story two widens story one. Routing is not only model-to-model. Routing is model, provider, identity, billing, region, compliance, support, and fallback. If you only name the model, you have named the headline. You have not named the system.
Story three. Deadlines and standards have to become evidence. This is the less glamorous part of the episode, which means it is probably where the real work lives. OpenAI's April tenth response to the Axios developer tool compromise includes a clear May eighth, twenty twenty-six date.
OpenAI says that effective May eighth, older versions of impacted macOS desktop apps will no longer receive updates or support, and may not be functional. The affected list includes ChatGPT Desktop, Codex App, Codex C L I, and Atlas.
OpenAI says the incident involved a compromised Axios version used in a GitHub Actions workflow connected to macOS app signing. OpenAI also says it found no evidence that products or user data were compromised or exposed, and no evidence that the signing material was misused.
This is exactly the kind of story that organizations handle badly because it feels like it belongs to "security" until the deadline arrives. But endpoint readiness is not just a security issue. It is a workflow continuity issue.
If your team depends on a desktop app, a coding app, a command-line tool, or a browser-connected workflow, then "older versions may not be functional" is not a footnote. It is a dated operating condition. And dated operating conditions need owners.
Not awareness. Owners. Awareness is when someone says, "Yeah, we saw the post." Ownership is when someone can say: we have three hundred fourteen installs, two hundred eighty-two are updated, twenty-one are scheduled, eleven are unmanaged, the escalation owner is named, and the next checkpoint is tomorrow at nine.
One of those sentences makes work happen. The other one is a scented candle. This is where NIST belongs in the episode. NIST's AI R M F profile concept note for trustworthy AI in critical infrastructure was created in April and updated April eighth.
The important point for this show is not that everyone should instantly produce a giant compliance program by Tuesday. Please do not. Tuesday has enough going on. The useful point is that NIST gives us a shape: trustworthy AI, critical infrastructure, lifecycles, supply chains, stakeholders, risk management, communication across teams.
For an operator, that can become one simple artifact. Owner. Evidence. Due date. That is it. Pick one critical AI-enabled workflow. Name the owner. Name the risk being controlled. Name the evidence link. Name the due date. Name the reviewer.
Name the next checkpoint. Do not try to solve every policy layer in one sitting. Build one evidence map that people can actually use. Because the failure mode in AI governance is not usually a shortage of beautiful principles.
It is a shortage of boring evidence at the moment someone asks, "Are we actually covered?" And "I believe so" is a phrase that has done terrible things to calendars. This is the third operating rule: Do not let standards stay abstract.
Turn one standard signal into one evidence artifact. That is the move. NIST becomes useful when it changes the meeting. If the meeting starts with: "What does trustworthy mean?" you may be there for a while. If the meeting starts with: "Who owns the workflow, what evidence proves the control, and what is due by Friday?"
now you are in business. So let's connect the three stories. GPT five point five says: the work surface is getting more capable. Anthropic and Amazon say: the infrastructure underneath that capability is getting bigger, more integrated, and more strategically concentrated.
The OpenAI remediation deadline and NIST concept note say: your execution evidence has to become more concrete. Put together, the operating question is: Can your organization route AI work intentionally while the underlying model, provider, endpoint, and standards environment changes?
That sounds like a mouthful because it is. But the practical version is short: Who owns the route? That is it. The route from user to model. The route from model to fallback. The route from task to approval.
The route from endpoint risk to remediation. The route from standard to evidence. The route from incident to rollback. If those routes are named, the organization can move. If those routes are not named, everyone waits for the most confident person to narrate reality.
And that person may be talented. But confidence is not a control. It is a lighting condition. This is where teams usually get the workflow wrong. They treat the model name as the decision. "We are using GPT five point five."
Or: "We are using Claude." Or: "We are keeping the old model for now." That sounds like a decision. It is usually only a label. The real decision is: what work is allowed to route through that model, under what conditions, with what evidence, with which fallback, and with what human review before the output becomes consequential.
A model name is not an operating policy. It is the thing on the door. The operating policy is what happens inside the room. This matters because organizations often make model decisions at the wrong altitude. They approve a tool.
They approve a vendor. They approve a license. They approve a tier. Then they assume the workflow underneath that approval will behave itself out of respect. It will not. Workflows are not polite. They are hungry. Give a workflow a faster tool and it will use the faster tool.
Give it a stronger model and it will route more work there. Give it no fallback rule and people will invent one. Give it no evidence requirement and the output will start sounding official before anyone can explain why.
That is not because people are reckless. It is because people are trying to get the work done. And when the system does not provide a route, the operator creates a shortcut. This is the part I think leaders underestimate.
Most AI drift is not dramatic. It does not arrive with someone saying, "I am now violating the governance framework." It arrives as: "I used the better model because the deadline was tight." "I used the fallback because the new one was unavailable."
"I skipped the citation pass because the answer looked clean." "I let it draft the whole thing because the first paragraph was good." "I put it into the client deck because we needed something visual." Tiny decisions. Reasonable decisions.
Human decisions. Then, two weeks later, nobody can explain the actual workflow. The official policy says one thing. The logs say something else. The team remembers a third thing. The customer experience reflects a fourth thing. And now everybody is standing around an operational smoothie asking which ingredient made it weird.
That is the problem with unowned routes. You do not know which change mattered. Was it the model? The fallback? The prompt? The tool permission? The reviewer? The endpoint version? The provider path? The missing citation requirement? The team that got access earlier than everyone else?
If you cannot answer that, you cannot manage the change. You can only react to the symptoms. So I want to name four routing mistakes. Mistake one: approving at the app level and managing at the workflow level.
The app is approved, but the workflow is not classified. That means a low-risk internal summary and a high-risk customer-facing answer may both pass through the same tool with the same casual confidence. The control should sit at the workflow level.
What is the task? Who sees the output? What harm happens if it is wrong? What evidence does it need? Who approves it? Mistake two: treating fallback as a technical detail. Fallback is not just what happens when the primary model is unavailable.
Fallback is a quality condition. If the fallback is weaker at citations, the reviewer needs to know. If the fallback is slower, the service-level expectation changes. If the fallback is cheaper but more error-prone, the savings are not free.
If the fallback is different across teams, the organization may be producing different answers while believing it has one process. Mistake three: measuring usage instead of correction load. Usage tells you the tool is being used. Correction load tells you what the tool is costing humans after the first answer appears.
How much editing? How many missing sources? How many escalations? How much reviewer fatigue? How many tasks have to be rerun? How often does the better model reduce work versus move work into a harder-to-see review layer? That is the number people miss.
Mistake four: letting remediation deadlines live outside the model conversation. The May eighth macOS cutoff is not the same category as GPT five point five. But in a real team, they collide. The same people who want to use the new model may also be using desktop apps, coding tools, command-line tools, browser flows, local files, and developer environments.
If the endpoint tail is not patched, the upgrade conversation is incomplete. You do not get to say, "We are modernizing the workflow," while the actual machines running that workflow are wandering toward a cutoff date with no owner.
That is the uncomfortable but useful management sentence. Modern AI operations are not just model operations. They are route operations. Model route. Tool route. Data route. Endpoint route. Approval route. Fallback route. Evidence route. All of those routes can change independently.
And if you only watch the model announcement, you miss the actual operating surface. This is why the episode is called Model Routing Check, not New Model Celebration Hour. Though, to be fair, New Model Celebration Hour would probably have better catering.
But the routing check is the thing that makes the celebration survivable. It lets you say: yes, we can use the stronger model here. No, we do not use it there yet. If it is unavailable, fallback goes here.
If fallback is used, the output gets this label. If citations are missing, the answer cannot ship. If a workflow touches customer commitments, a human approves it. If the endpoint is not remediated, that machine is out of the path.
If the standard asks for evidence, the evidence lives in this artifact. That is how a model upgrade becomes a controlled change. Not slower. Cleaner. Here is the Monday action block. Forty-five minutes. One owner. No mythology. Minute zero to eight: pick one AI workflow where model quality matters.
Not the whole company. One workflow. Support escalation. Legal summary. Sales proposal. Engineering code review. Research memo. Financial analysis. Name the primary model and the fallback model. If nobody knows the fallback, congratulations, you found the first action item.
Minute eight to sixteen: run the top ten prompts across the primary and fallback paths. Do not just ask if the new output is better. Measure: material differences, missing citations, review time, correction load, latency, and any output that changes the decision a human would make.
Minute sixteen to twenty-four: write the routing rule. Use the stronger model for these workflows. Use fallback for these conditions. Require human approval here. Require citation here. Block autonomous action here. Escalate to this owner when the route changes.
Minute twenty-four to thirty: add the access matrix. Team. Plan or tier. Primary model. Fallback model. Known limits. Last checked date. This is not fancy. That is why it works. Minute thirty to thirty-seven: tie in the live remediation deadline.
For the OpenAI macOS app cutoff, name the owner, affected products, version floor, remaining installs, and next checkpoint. If it is not relevant to your team, pick the equivalent dated dependency in your environment. The point is to practice dated ownership, not to worship one vendor deadline.
Minute thirty-seven to forty-five: create one evidence map. Workflow. Risk. Owner. Control. Evidence link. Due date. Reviewer. Next checkpoint. One page. If it becomes a forty-slide deck, something has gone wrong and possibly someone needs water. That is the model-routing check.
Primary. Fallback. Access. Prompt regression. Correction load. Remediation owner. Evidence map. It sounds small because it is supposed to be runnable. The whole point of this show is that control has to survive contact with a normal work week.
Not the strategy offsite. Not the executive memo. The normal week. The week with meetings, shipping pressure, weird access differences, one person out sick, one person overconfident, and one vendor update that lands while everyone is pretending Tuesday is under control.
And I want to be careful about the tone here. This is not an anti-upgrade episode. Better models are useful. Faster work is useful. Deeper cloud integrations can be useful. Standards work is useful. Security remediation is useful.
The issue is not progress. The issue is unowned progress. Unowned progress is how organizations get the benefits loudly and the risks quietly. And quiet risk is the most expensive kind because nobody budgets for it until it has a name, a meeting, and a story people wish they had prevented.
So the closing question is this: Which AI workflow in your organization has already changed routes, even though nobody has formally approved the new route? Not which model is best. Which route is real. Who gets access. What fallback happens.
Who reviews it. What evidence proves it. What deadline is attached. And who can stop it if the route is wrong. That is the control. The model can get smarter. The infrastructure can get larger. The deadline can get closer.
The standard can get more serious. But the organization only gets better if the route is owned. __OUTRO_MUSIC__ That is AI Change Desk for today. The practical next step: run the forty-five minute model-routing check on one workflow before the new model path becomes default.
I will put the sources, worksheet, and show notes in the episode page. As always: watch the route, name the owner, and do not let "we upgraded" substitute for "we controlled the change."