Production script transcript
Preview Before Power Mode
EP036 · Jun 29, 2026 · 16m 39s
Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.
Imagine the most powerful model in the building is not on every desk. Not because procurement forgot the form. Not because I T is hiding in a server closet with a clipboard. Although, spiritually, sometimes yes. It is limited on purpose.
A few trusted groups. A few approved surfaces. A P I access here. Codex workspace access there. ChatGPT, not yet. No public waitlist. No everybody-clicks-the-big-shiny-button moment. And that is where the real episode starts. Because most organizations hear frontier model preview and immediately start acting like it is a weather event.
It is coming. We should prepare. Somebody buy umbrellas. Maybe the umbrellas need admin rights. But this is not weather. This is access. And access is a management decision. Who gets the strongest capability? In which workspace? With which data?
With which tools? Under which safeguards? With what evidence? And who can roll it back when the preview stops being a preview and starts becoming dependency? That is the shift today. The question is not whether the model is impressive.
The question is whether the organization has a gate strong enough for the capability behind it. Before power mode becomes normal mode, name the preview gate. Welcome back to AI Change Desk. I am Michael. Today is Monday, June twenty-ninth, twenty twenty-six.
And this is episode thirty-six: Preview Before Power Mode. The lead signal is OpenAI's June twenty-sixth preview of GPT five point six Sol. Sol, as in sun. Which is a little on the nose. Because apparently we are now naming frontier models like someone opened a mythology cabinet and a weather app at the same time.
The naming is not the point. The operating pattern is. OpenAI says it is beginning a limited preview of the GPT five point six family. Sol is the flagship model. Terra is the balanced model. Luna is the faster, lower-cost model.
The announcement frames Sol as a stronger frontier system, especially around software engineering, computer use, professional knowledge work, scientific research, and cybersecurity. And then comes the sentence operators should care about. The preview is limited. It starts with trusted partners and organizations.
Broader availability is planned later. Stronger capability arrives with stronger safeguards. That is the story. Not the benchmark chart. Not the model-family naming scheme. Not the internet doing its usual thing where half the room says this changes everything and the other half says actually my spreadsheet still hates me.
The story is that frontier capability is being treated as scoped power. And scoped power needs governance before it becomes habit. Preview access is not a perk. It is a control surface. The companion OpenAI help article makes the gate much more concrete.
During the preview, GPT five point six is available through the OpenAI A P I and Codex to a limited group of trusted partners and organizations. It is not available in ChatGPT during the preview. The preview is not a broad self-service program.
It is not available to individual consumers. There is no public application or waitlist. And the access boundary is surface-specific. Approval for the A P I does not automatically include Codex. Approval for Codex does not automatically include the A P I.
That detail matters. Because a lot of organizations still treat AI access like one big permission bucket. Approved. Not approved. Pilot. Blocked. Maybe approved if the vice president forwards a vendor webinar with the subject line, thoughts? But frontier access does not work cleanly in one bucket.
A model in an A P I can become an application dependency. A model in Codex can become a code-change dependency. A model in ChatGPT can become a knowledge-work dependency. A model in a spreadsheet can become a decision dependency.
Those are not the same risk. They are not the same evidence trail. They are not the same rollback path. So the first operating lesson is simple. Do not approve the model. Approve the surface. Who can use it in the A P I?
Who can use it in Codex? Who can use it in internal tools? Who can connect it to sensitive data? Who can let it write, execute, commit, or publish? And who can turn it off? That last question keeps showing up on this show because apparently the off switch remains civilization's most underrated technology.
The wheel was good. Fire had a strong run. But a named off switch? Chef's kiss. The second lesson is that stronger safeguards are not a reason to stop managing. They are a reason to manage more clearly.
OpenAI describes layered safeguards, including model-level protections and real-time checks. The help article says some requests may be blocked or take longer while additional safety checks run, especially in dual-use areas like biological and cybersecurity work. That is useful.
It is also operationally messy. Because when a request takes longer, or gets blocked, or needs a narrower scope, someone has to know what happens next. Is the user supposed to retry? Escalate? Switch models? Rewrite the task?
Ask security? Ask legal? Ask the person who originally said, sure, let's pilot this, and then vanished into a steering committee? A safeguard is not a workflow. A blocked response is not an incident process. A delay is not a review queue.
If the preview touches real work, your organization needs routing. What requests are allowed? What requests require additional review? What requests are out of bounds? What gets logged? Who sees the log? What does the user do when the model says no, waits, or returns less than expected?
Because if the answer is everyone improvises in Slack, congratulations. You have invented governance jazz. It may be expressive. It is not reliable. The third lesson comes from OpenAI's June twenty-fifth post on how agents are transforming work.
That post is important because it explains why this is not just a model-access story. OpenAI says Codex work is getting longer, more embedded, and more cross-functional. It describes people using Codex for work that would take a human more than thirty minutes, more than one hour, and sometimes much longer.
It also describes internal adoption spreading beyond engineering into legal, finance, recruiting, and other functions. Keep the attribution clear. That is OpenAI reporting on OpenAI and its users. It is not neutral market proof. But it is still useful directional evidence.
Because the pattern is familiar. As tools get more capable, people stop using them for little questions. They start delegating work. Then they start delegating parallel work. Then someone says the phrase, autonomous workflow, in a meeting, and suddenly half the room is smiling and the other half is quietly updating the risk register.
That is why the Sol preview matters operationally. A stronger model does not enter a vacuum. It enters an environment where agent work is already getting longer. It enters codebases. It enters workspaces. It enters budget lines. It enters review queues.
It enters the weird part of the organization where everyone agrees something is important and nobody owns the handoff. Stronger models amplify old ownership gaps. They do not magically clean them up. Microsoft's June twenty-fifth Copilot in Excel announcement gives the cross-vendor version of the same point.
This is not a frontier model announcement. It is a high-stakes work-surface announcement. Microsoft is putting more Copilot capability into Excel for finance workflows. The useful part is not just that Copilot can help in a spreadsheet. The useful part is the structure around the work.
Skills for repeatable workflows. Trusted data connectors. Planning. Traceability. The ability to show work instead of just producing an answer with the confidence of a consultant who has never seen the workbook before. Finance is a good stress test because finance people have a beautiful intolerance for mysterious numbers.
You cannot walk into a close meeting and say, the model vibe is forty-seven million. People ask follow-up questions. Cruelly specific ones. Where did the number come from? Which data source? Which assumption? Which formula? Who changed it?
When? Can we reproduce it? Can we explain it to an auditor without everyone suddenly discovering a dentist appointment? That is the governance lesson. As AI gets more powerful, the interface has to become more accountable. Plans matter.
Skills matter. Data connectors matter. Change traces matter. Reviewable steps matter. Because in high-stakes work, the magic answer is not enough. The work has to leave receipts. Power without trace is just confidence with a loading spinner. Anthropic's Claude Tag gives the shared-channel version.
Claude Tag starts in Slack. Teams can grant Claude access to selected channels and selected tools. People in the channel can tag Claude and delegate work. Anthropic says Claude builds context from the channels it is in, can remember relevant information, can work asynchronously, and can plan tasks into the future.
The control details matter here too. Administrators define what information and tools Claude can access in which channels. Memory stays scoped to the channels defined by administrators. Administrators can set spend limits. They can view a log of what Claude has done and who requested each task.
That is the same pattern again. Stronger or more proactive AI work needs a boundary. Not a vibes boundary. Not a please be careful boundary. An actual boundary. Which channel? Which tools? Which memory? Which spend limit? Which logs?
Which owner? If a shared agent can remember the channel, work in the background, and follow up over time, then the channel is no longer just a conversation space. It is an operating surface. And operating surfaces need controls.
Otherwise, you do not have a teammate. You have a very polite raccoon with channel history. And the raccoon has a budget. So what should operators do today? Run a forty-five minute Preview Before Power Mode review. Not a model admiration session.
Not a vendor-roadmap séance. A gate review. Bring security, I T, legal, procurement, the business owner, and the person who will actually get yelled at if the workflow breaks. That last person is often the only adult in the room.
Answer eight questions. First: Who is eligible? Name the group. Not everyone interested. Not anyone with a strong opinion and a hoodie. A named group. Trusted users, approved team, pilot cohort, whatever fits your organization. But name it.
Second: Which surface is approved? A P I? Codex? Internal app? Spreadsheet? Chat interface? A channel agent? Do not let one approval quietly become all surfaces. Third: What data can it touch? Public data. Internal non-sensitive data. Customer data.
Security data. Financial data. Code. Production logs. If the data class is vague, the approval is vague. Fourth: What tools can it use? Read-only tools are one thing. Write tools are another. Code execution is another. Commit access is another.
Deployment access is a different animal entirely. That animal has teeth. Fifth: What safeguard routing exists? If the model blocks a request, delays a request, or asks for narrower scope, what happens? Does the user retry? Escalate? Open a ticket?
Switch models? Stop? If nobody knows, write the path before the preview expands. Sixth: What is the spend and cache budget? Stronger models change the cost profile. Longer context changes the cost profile. Prompt caching can help, but it is still part of the budget design.
Name the owner before the invoice becomes a governance document written in dollar signs. Seventh: What evidence is retained? Prompt. Plan. Tool calls. Data sources. Outputs. Approvals. Review notes. Exceptions. Rollback decisions. The preview is not real governance unless the work can be reconstructed.
Eighth: Who can remove access? Not just who can grant it. Who can narrow it? Pause it? Revoke it? Expire it? Move it from preview to blocked if the evidence is bad? That is the whole test. Eligible group.
Approved surface. Data boundary. Tool boundary. Safeguard routing. Spend and cache budget. Evidence log. Rollback owner. If you cannot answer those eight, you are not ready for power mode. You are ready for a demo. Demos are fine.
Just do not confuse the demo with the operating model. The deeper pattern is bigger than OpenAI. This is not a Sol episode. It is a preview-governance episode. OpenAI shows the frontier access gate. Microsoft shows the traceable high-stakes work surface.
Anthropic shows the shared-channel agent boundary. Episode thirty-four showed what happens when AI-assisted work moves toward production security fixes. All of those point to the same management discipline. Capability does not become safe because it is impressive. It becomes usable when the organization can say: who gets it, where it runs, what it touches, what it costs, what it leaves behind, and who can stop it.
That is boring in the best possible way. Boring is how mature controls feel from the outside. The exciting version is the incident review. Nobody should aspire to make governance exciting in that specific way. So here is the closing thought.
When the strongest model is still behind a preview gate, do not treat that as an inconvenience. Treat it as a rehearsal. This is the moment to practice scoped access before the tools become normal. This is the moment to separate A P I approval from workspace approval.
This is the moment to decide which data is allowed, which tools are allowed, which logs matter, and who owns the rollback. Because once frontier capability becomes everyday workflow, the organization will not suddenly become more disciplined. It will become faster.
And speed without ownership is how small approval gaps turn into expensive archaeology. Preview before power mode. Name the gate before the gate disappears. That is AI Change Desk for today. If this helped, share it with the person who owns access approvals before the next model rollout turns into a scavenger hunt.
I am Michael. I will see you next time.