Skip to content
MHBMMichael Hanna-Butros MeyeringComplex systems · human outcomes
Menu

Full transcript

Who Approved the Question?

EP045 · Sep 14, 2026 · 21m 20s

Source record

Sources cited in this episode

Public references used to prepare the companion episode.

  1. OpenAI, Now everyone can put data to work, September 10, 2026
  2. OpenAI Help Center, Using the Data plugin in ChatGPT Work and Codex, updated September 10, 2026
  3. Google Workspace Updates, Context-aware access controls are available for Gemini Enterprise in the Admin console, September 8, 2026
  4. Google Workspace Updates, Manage external sharing for Gemini Notebook in the Admin console, September 10, 2026
  5. NIST, Privacy Framework: Getting Started, updated August 29, 2025

Someone types one sentence. Show me which customers are becoming expensive to serve. No query language. No ticket to the data team. No meeting where twelve people debate the meaning of one column and somehow schedule another meeting.

The A I checks approved sources. It uses the person's existing permissions. It compares the numbers. It builds a dashboard. It recommends who should act. And now one ordinary sentence has become a data product. The user may be allowed to see every row.

That does not answer whether this was the right question. The right grouping. The right metric. The right audience. Or the right decision to make from the answer. Natural language is wonderful. It also lets bad analysis wear business casual.

Permission can open the door. It does not write the purpose. The next governance gap is not whether the A I can reach the data. It is whether anybody approved what the question turns into. Quick disclosure before we start.

AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.

One additional disclosure as the Desk evolves. My current role includes privacy work, so you will hear me pay closer attention to purpose, access, retention, deletion, and accountability. I will not discuss nonpublic work here. These are my personal views, and they do not represent the State of Oregon or any other organization.

Opening music

Welcome back to AI Change Desk. I am Michael. This is episode forty-five. Who Approved the Question? Today is Monday, September fourteenth, twenty twenty-six. And as I checked the sources this morning, the most useful signal was not another model score.

It was not a leaderboard. It was not a demo where the cursor moves by itself and everybody applauds like the spreadsheet has achieved consciousness. It was a product change that makes a normal business question much more powerful.

On September tenth, OpenAI introduced the Data agent in ChatGPT Work. OpenAI says the Data agent can connect to approved company sources, investigate what changed, build interactive dashboards, share findings, and help carry out approved actions through connected tools.

The sources can include data warehouses, business intelligence systems, Google Drive, SharePoint, and other connected services. The pitch is straightforward. People should be able to ask important business questions without waiting for someone else to translate every question into a query.

That can be genuinely useful. It can also move the control surface. Because when the cost of asking falls, the number of people asking goes up. The number of variations goes up. The number of dashboards goes up.

And the distance between curiosity and action gets shorter. OpenAI says administrators choose which data connections are available and which roles may use them. It says queries enforce the connected account's existing permissions, including table, row, and column restrictions.

That is an important control. It is also where the conversation can stop too early. Because existing permission answers one question. Can this identity retrieve this information? It does not automatically answer the next six. Why is this person asking?

What definition is the system using? What information is missing? Which groups are being compared? Who will receive the result? And what decision will the result influence? Access control answers who can retrieve. Purpose control answers what they may ask the data to do.

Those controls can support each other. They are not interchangeable. This is the next step in the sequence we have been building. In episode forty-one, we asked what the A I could see. In episode forty-two, we followed the retention boundary.

In episode forty-three, we asked whether the safeguard actually covered the run. In episode forty-four, we treated the audit trail as data that needed its own owner. Now we have to govern what people ask the system to conclude.

Not because questions are dangerous by default. Because questions become processing instructions. And with an agent, the processing instruction can become a dashboard, a message, a recommendation, or an action before the organization realizes it crossed from exploration into operations.

Start with the question itself. Show me which customers are becoming expensive to serve. That sounds useful. What does expensive mean? Support tickets? Refunds? Cloud usage? Late payments? Manual exceptions? Time spent by account teams? And what does customer mean?

An account? A household? A contract? A person? Those are not formatting details. They determine who appears in the answer and why. OpenAI's guidance recommends a semantic layer. That means authoritative business definitions, custom calculations, and relationships between datasets that help keep the analysis on track.

That is the right direction. But a semantic layer is not just technical plumbing. A metric definition is policy expressed as math. It decides what counts. It decides what does not count. It decides when the clock starts.

It decides which denominator makes the result look calm, and which denominator makes the room suddenly interested. The metric had one job. Then three departments gave it three definitions and invited it to a steering committee. If an A I can build analysis faster, the organization needs a faster way to prove which definition it used.

Not the definition someone remembers. The version that was active for that run. The calculation. The filters. The comparison period. The exclusions. The joins. And the owner who says, yes, this is the meaning we intended. OpenAI's own help guidance tells users to check the source, time period, filters, and metric definition before relying on a result.

That line matters. Because a polished dashboard can make an unsettled definition look finished. A dashboard is a confident rectangle. Confidence is not provenance. If the analysis differs from an existing report, the documented advice is to compare the sources and definitions.

That should not be a troubleshooting trick. It should be part of the receipt before the result influences a real decision. What source was used? How current was it? Which records were unavailable? Which filters were active? What did the system infer?

What evidence contradicted the answer? Who reviewed the difference? And what happened next? There is another detail in the help documentation that operators should notice. OpenAI says the Data plugin can be triggered implicitly. You may not always need to mention it by name.

The guidance says that if you are unsure, you can ask whether the plugin was used. That is convenient. It also means provenance should not depend on whether a user remembers to ask after the fact. The run should tell you.

Which plugin was invoked? Which connected account did it use? Which source did it query? Which definition did it apply? And which parts came from the model rather than the source? If the system can move from ordinary conversation into enterprise analysis, the receipt needs to move with it.

The user should not have to perform conversational archaeology to discover what happened. Then comes publishing. OpenAI says that when an analysis is published through OpenAI Sites, the data used in the analysis is copied into the published site.

Read that carefully. The source permission controlled the query. The published artifact creates a copy. The copy has an audience. The copy has a location. The copy may have a refresh cycle. The copy may remain useful after the source data changes.

And the copy can be perfectly accurate while still reaching the wrong people. This is why data-agent governance cannot end at connector approval. The connector answers whether the system can reach the source. The publication control answers whether this derivative result can leave the conversation.

The audience decision answers who may see it. The retention decision answers how long the copy remains. The refresh decision answers whether tomorrow's data silently changes yesterday's approved artifact. Different gates. Different owners. One workflow. Google's recent updates make that separation visible from another direction.

On September eighth, Google announced context-aware access controls for Gemini Enterprise. For eligible customers, administrators can apply conditions such as device security and location at an organizational-unit or group level. That can help answer whether a user may enter the A I surface from this device, in this context, under this identity.

Useful. Still not a purpose decision. On September tenth, Google also announced more granular external-sharing controls for Gemini Notebook. The documented options include off, trusted domains, any external email address, and public-link sharing. Google says external sharing is off by default and can be controlled by domain, organizational unit, or group.

Again, useful. And again, the settings expose separate questions. Can the user enter? Can the A I retrieve? Can the artifact be shared? Can it be shared externally? Can it be public? Should this particular result be shared with this particular audience?

The first five can be settings. The last one is a decision. A permissive setting is not a standing order to use the maximum permission. That principle matters in privacy work because privacy risk is not only a breach problem.

NIST's voluntary Privacy Framework describes privacy risk as something that can arise from normal system operations with data across a full lifecycle. Collection. Use. Sharing. Retention. And disposal. The system can function exactly as designed and still create a problem for a person or group.

That is why security and privacy overlap, but they are not the same job. The account can be secure. The query can be authorized. The dashboard can be accurate. And the proposed use can still deserve a no.

Or a narrower question. Or a different audience. Or a human review before action. This is operational guidance, not a universal legal conclusion. The right legal analysis depends on jurisdiction, data, role, purpose, contract, and decision context. But operators do not need to wait for a lawsuit to ask better questions.

They need a receipt. And this does not mean every question needs a committee. That would defeat the point. Use three lanes. The first lane is exploratory. The user is trying to understand the data. The result stays with a small approved audience.

It is labeled as exploratory. It does not trigger a consequential action. The second lane is operational. The question is recurring. The definitions are approved. The sources and audience are known. The output supports routine work under a named owner.

This is where a reusable receipt can make the workflow fast. The third lane is consequential. The result may affect a person's access, opportunity, service, priority, review, or treatment. That lane deserves stronger evidence, meaningful human review, and a correction path before action.

The names of the lanes can change. The separation should not. Do not force a brainstorming question through the same process as a decision that affects someone. And do not let a consequential decision borrow the casual controls of brainstorming because both started in a chat box.

Scale the review to the consequence, not to how easy the interface feels. Here is the Question-to-Action Receipt. Seven parts. First, the purpose receipt. Write the exact business question. Name the allowed use. Name the prohibited use. Name the outcome owner.

If the purpose is just find something interesting, the result stays exploratory. It does not become an operating decision because the chart has rounded corners. Second, the identity and authority receipt. Who asked? Which agent ran? Which connected account authenticated?

Which role and source permissions were effective? Was the access personal, shared, or agent-owned? Existing permission belongs in the receipt. It just does not finish the receipt. Third, the source receipt. List the systems, datasets, documents, classifications, freshness, and known exclusions.

What was missing? What was delayed? What was copied from a report rather than queried from the source? An answer built from incomplete data can be useful. It must not pretend the missing data voted yes. Fourth, the meaning receipt.

Record the metric definitions, calculations, joins, cohorts, comparison periods, and semantic-layer version. If two teams use the same word for different math, the A I did not resolve the disagreement. It automated it. Fifth, the evidence receipt. Keep the method, filters, uncertainty, contradictions, validation, and reviewer.

What would change the conclusion? What did the system not test? What did the human actually review? Approved is not useful if nobody can explain what was approved. Sixth, the audience and copy receipt. Who receives the answer?

What data moves into the dashboard, site, email, slide, or message? Where does that copy live? Who can reshare it? How long does it remain? Does it refresh automatically? If the artifact creates a new copy, give the copy its own boundary.

Seventh, the action and disposition receipt. What did the system recommend? Who approved the next step? What happened downstream? Who could be affected? How can a bad conclusion be corrected? And what was the final disposition? Accepted. Narrowed.

Held. Rejected. Reversed. Still under review. The analysis ended is not a disposition. Here is how the boundary moves during a completely normal week. Monday, a service team asks the A I to explain why support costs increased.

The question is legitimate. The agent uses approved ticket data, account data, and the team's standard cost definition. It finds that a small group of accounts is generating more manual work. So far, so useful. Tuesday, someone asks for a breakdown by account manager and region.

That may also be reasonable. But the analysis now describes employee activity as well as customer behavior. The purpose changed a little. The audience may need to change with it. Wednesday, the dashboard is published for leadership. The source permissions still worked exactly as designed.

But the published artifact is now a separate copy, and more people can see the combined result than could query every source directly. Thursday, a manager asks the system to rank the account managers most responsible for the increase.

The dashboard can probably produce a ranking. That does not mean the ranking is the right use of the data. Ticket volume may reflect account complexity, product defects, regional staffing, or a manager who documents work more carefully than everyone else.

The data may answer the calculation while failing the question. Friday, the ranking appears in a meeting as if it had always been the purpose. Nobody changed a permission. Nobody broke into a database. Nobody typed please create an unfair conclusion.

The workflow simply moved from cost analysis, to employee comparison, to leadership distribution, to a consequential interpretation. One reasonable follow-up at a time. Purpose creep rarely arrives wearing a cape. Usually it arrives as while we are here, can we also add one more column?

With a Question-to-Action Receipt, Monday's analysis can proceed under the approved service-cost purpose. Tuesday's new grouping triggers an audience and use review. Wednesday's publication records the copied fields and recipients. Thursday's ranking is held until the team tests whether the metric supports that interpretation.

And Friday's meeting receives either a bounded result, an explicit warning, or no ranking at all. That is not bureaucracy winning. That is the organization keeping a useful analysis from becoming a confident mistake. You can test this in forty-five minutes.

Pick one recurring business question that an A I can help answer. Not the whole analytics estate. One question. For the first seven minutes, write the allowed purpose, the prohibited use, and the owner. For the next seven, identify the requester, agent, connected account, and effective permissions.

Then spend seven minutes on sources. List classifications, freshness, known exclusions, and missing records. Spend seven minutes locking the meaning. Metric. Calculation. Join. Cohort. Comparison period. Version. Use the next seven minutes to inspect evidence and review. Source.

Filter. Uncertainty. Contradiction. Reviewer. Then use six minutes for the audience and copy. Who sees it? What gets copied? Where does it live? When does it refresh or disappear? Use the final four minutes to decide. Approve. Narrow.

Hold. Or stop. Write the owner and the correction path. The goal is not to make every question slow. The goal is to make repeated questions governable. Once the purpose, definitions, sources, audience, and review pattern are proven, the organization can reuse them.

That is how governance increases speed instead of becoming the person standing in front of the copier with a clipboard. Track a small scorecard for one month. What percentage of recurring A I analyses have an approved purpose?

What percentage use a named semantic definition? What percentage record source, period, filters, and exclusions? What percentage of published artifacts have a named audience and retention rule? What percentage of consequential recommendations receive meaningful human review? How often can the team reconstruct why a result changed?

And how quickly can it correct a conclusion that should not have become action? Those are operating measures. Not dashboard furniture. The headline this week is that more people can ask enterprise data harder questions in plain language.

That is a capability gain. The governance response should not be panic. It should not be a blanket ban. And it should not be the permissions already handle it. The response is to separate the gates. Access. Purpose.

Meaning. Evidence. Audience. Copy. Action. Do not scale the answer until you can govern the question. That is the Desk for this week. If this helped, share it with the person who owns data governance, privacy, analytics, or the increasingly brave soul responsible for all four.

The companion field guide and Question-to-Action Receipt will be linked with the episode. Until next time, keep the question visible, keep the evidence attached, and keep the final decision human-owned.

Closing music