Full transcript
Who Owns the AI Audit Trail?
EP044 · Sep 7, 2026 · 20m 55s
Source record
Sources cited in this episode
Public references used to prepare the companion episode.
- Google Workspace Updates: Introducing comprehensive audit logs for Gemini Notebook in the Workspace Admin console
- Google Workspace Admin Help: Gemini Notebook log events
- Google Workspace Admin Help: Gemini Notebook BigQuery log schema
- Anthropic: Developing Enterprise Frontier Safeguards with our customers
- OpenAI: Offering Zero Data Retention for frontier models
Good news. The AI notebook finally has an audit trail. You can see who did what. You can see when they did it. You can see what they shared. You can see the source they used. You can see the artifact they made.
Now the less comfortable news. The audit trail is another dataset. It can carry a person's identity. An I P address. A source name. A source U R L. The visibility of a notebook. The identifier for an infographic, an audio overview, or a report.
And if you export it, congratulations. You now have the second dataset in a second system, managed by a second group of people, under a retention rule nobody remembers approving. Governance has a remarkable ability to create more things that need governance.
That does not mean we stop logging. It means we stop pretending the log lives outside the privacy map. Because the prompt is data. The output is data. The source is data. And the receipt proving what happened?
That is data too. Quick disclosure before we start. AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice.
These are my opinions and are not representative of any organization. One additional disclosure as the Desk evolves. My current role includes privacy work, so you will hear me pay closer attention to purpose, access, retention, deletion, and accountability.
I will not discuss nonpublic work here. These are my personal views, and they do not represent the State of Oregon or any other organization.
Opening music
Welcome back to AI Change Desk. I am Michael. This is episode forty-four. Who Owns the AI Audit Trail? Today is Monday, September seventh, twenty twenty-six. It is Labor Day here in the United States. If you are off today, I hope you get some real time away from work.
If you are working, thank you. And the date matters to this conversation for a simple reason. Systems built to monitor work can also create records about workers. That does not make monitoring wrong. It makes purpose, transparency, access, and restraint part of the design.
As I source-checked this on Sunday, September sixth, the newest concrete signal came from Google Workspace. Google announced comprehensive audit logs for Gemini Notebook. The headline sounds administrative. The operating consequence is much bigger. The organization can now see more of the control trail around notebook activity.
And that control trail has its own privacy architecture. The operating question is this. Can your organization prove what its AI control systems record, where those records live, who can read and export them, how long they persist, and how they reconnect to people and source content?
We mapped the content. We mapped the retention. We tested the safeguard. Now we have to govern the receipt. On September third, Google announced a gradual rollout of Gemini Notebook log events for eligible Workspace customers with access to its audit and investigation tools.
Google says administrators can review how users interact with Gemini Notebook to generate and summarize content. The documented fields are not vague. They can include the actor's email address. The date and time. The event. An I P address.
An I P network number and region code. The visibility of the notebook before and after a change. The resource identifier, title, type, and owner. The source identifier. The source name. The source type. The source U R L.
And identifiers for Studio artifacts such as an audio overview, infographic, report, slide deck, or video overview. That is useful evidence. It is also a useful reminder that metadata is not decorative exhaust. A source name can reveal the subject of the work.
A source U R L can reveal where the work came from. A visibility change can reveal who opened the door. An actor field can connect the event to a person. An I P address can add network context and sometimes approximate location, although Google's own documentation warns that it may instead reflect a proxy or a V P N.
This is not a reason to panic. It is a reason to map the fields before turning on every export because the button looks lonely. Google says administrators can investigate the events in Workspace tools. It also documents an optional export path into Big Query.
That export is disabled until an administrator enables it. That detail matters. Before the export is enabled, the organization has one control record in one governed service. After the export is enabled, it has a copy in a data warehouse, with warehouse permissions, warehouse retention, warehouse queries, warehouse service accounts, and perhaps a very enthusiastic dashboard named something like notebook governance final version two, actually final.
Let's be honest. The export may be completely justified. It may support investigations, trend analysis, compliance evidence, or incident response. But the reason needs to exist before the copy. Otherwise, the organization is not building evidence on purpose.
It is collecting first and inventing the purpose during the audit. There is another detail in Google's announcement that deserves its own line. Google says audit log storage follows standard Workspace regional routing policies. But Gemini Notebook user data, including notebooks, sources, and chat histories, is stored globally and does not currently support data regionalization.
Those are two different statements about two different planes. The regional path of the audit record does not prove the regional path of the underlying content. And the global path of the content does not tell you every detail about the log.
If a diagram has one box labeled Gemini Notebook, it is probably too simple. Think of the workflow as two parallel planes. The first is the content plane. That is the notebook. The sources. The copied text. The local uploads.
The links. The conversations. The generated audio, report, slide deck, infographic, quiz, or video. The second is the control plane. That is the actor record. The event. The visibility change. The source identifier. The alert. The investigation. The reviewer decision.
And the export. The content plane tells you what the system worked with. The control plane tells you what the organization can prove about that work. Both can be sensitive. Both can be incomplete. Both can move. And both need an owner.
This is where the privacy question gets more interesting than "Do we retain prompts?" That question still matters. It is just no longer enough. A system can delete the prompt and keep a source name. It can hide a notebook from the user and retain a separate administrative record.
It can regionalize the log and globally store the content. It can keep content under one access policy and export metadata under another. It can restrict human review of the content while still sending a narrow safety signal somewhere else.
Each arrangement can be reasonable. But none of them is explained by one word like private, secure, logged, or zero retention. The architecture is the answer. And the receipt has to show the architecture that was actually active.
This same pattern is appearing in other frontier A I systems. Anthropic announced Enterprise Frontier Safeguards on September first. Anthropic describes it as a design that combines zero-data-retention access with monitoring across interactions. Its stated approach lets activity data used for monitoring remain in cloud infrastructure controlled by the customer, under the customer's encryption keys, access policies, and audit logging.
Anthropic says monitoring signals go to the customer for review, so the person reviewing a flag can be someone the customer has trained and cleared for sensitive material. That is an important architecture signal. It is not proof that every implementation is ready or effective.
Anthropic says the service will roll out in phases beginning later this fall. So treat it as an announced design and rollout commitment, not as a universal capability already running everywhere. OpenAI previewed a related but different approach in August.
Private Safety Processing is designed to look for patterns across related interactions without giving OpenAI personnel access to the underlying customer content. OpenAI says the content can remain in customer-controlled infrastructure, or in OpenAI-provided storage encrypted with keys controlled by the customer.
When the system identifies a risk, OpenAI says it receives a narrowly defined signal about the type of activity, not the underlying content. As of the source check for this episode, OpenAI described the system as being tested with early customers.
It said rollout and a technical paper were planned for September. Planned is not completed. Previewed is not generally available. And a design claim is not a runtime receipt. Still, the direction is clear enough to matter. Providers and customers are trying to separate content custody, automated safety analysis, human review, and enforcement signals.
That separation can improve privacy. It can also make responsibility harder to see if nobody maps the handoffs. Who holds the content? Who holds the key? Who receives the signal? Who can inspect the evidence? Who decides the flag is real?
Who can stop the workflow? Who retains the decision? And who can correct the record when the signal is wrong? That last question matters. A control plane should not only prove that an event happened. It should preserve the path from event to decision.
Otherwise, the log can tell you that someone was flagged without telling you whether the flag was valid, what happened next, or whether the record was ever corrected. A log without a disposition is an accusation-shaped data row.
That is not enough for governance. So here is the practical framework. I call it the Two-Plane Privacy Receipt. Seven parts. First, the purpose receipt. Why are you collecting each class of content and control evidence? What decision will it support?
What uses are outside the approved purpose? Evidence is a written purpose tied to an owner and a decision. The failure mode is collecting everything because storage is cheap and questions are expensive. The decision rule is simple.
If nobody can state the use before collection or export, do not collect or export it by default. Second, the content-plane receipt. List the sources, conversations, uploads, links, generated artifacts, sharing states, and sensitive categories the workflow can create or touch.
Evidence is a real sample workflow and its actual artifacts, not the vendor's generic architecture slide. The failure mode is mapping prompts and outputs while ignoring uploaded files, source links, copied text, generated media, or shared notebooks. The decision rule is that an unknown content path stays out of production.
Third, the control-plane receipt. List the log events, identities, network fields, source metadata, labels, alerts, reviewer decisions, and exports created around the workflow. Evidence is the current schema and a test event. Google's own help page says not every attribute appears for every event, and that the list is not exhaustive and may change.
So the failure mode is treating a screenshot from launch day as a permanent data dictionary. The decision rule is to review schema changes like product changes, because new evidence fields are new data fields. Fourth, the location receipt.
Where is content stored? Where are logs stored? Where is processing performed? What follows regional routing? What is global? Where do exports land? Where do backups and incident copies go? Evidence is a service-specific map with source dates and contract references.
The failure mode is using one location answer for two different data planes. The decision rule is that a regional audit record does not prove regional content storage. Fifth, the access receipt. Who can open the content? Who can search the logs?
Who can enable export? Who can query the warehouse? Which service accounts can move the data? Who can view content inside an investigation? Evidence is role assignments, privileges, service identities, and a readback from the real environment. The failure mode is giving the people who investigate misuse a permanent ability to browse everything, whether or not a case exists.
The decision rule is least privilege with a reason, a review date, and separation between routine administration and sensitive investigation. Sixth, the lifecycle receipt. When does each record appear? How much lag exists? How long does it remain available?
What happens when a user deletes the notebook? What happens under a hold? What happens to a Big Query export? What happens to backups? What happens when the schema changes? Evidence is retention and deletion testing across the original service and every downstream copy.
The failure mode is deleting the content and leaving the export forever. The decision rule is that every new copy inherits an explicit lifecycle before activation. Seventh, the correlation-and-action receipt. How can the control record reconnect to a person, a source, a notebook, or another system?
What alert or investigation can follow? Who reviews the result? What action can they take? How can an affected person correct an error where correction is appropriate? And what closes the case? Evidence is one safe test from event through disposition.
The failure mode is a beautiful warehouse full of evidence nobody can translate into a responsible decision. The decision rule is this. If a log cannot support a defined response, or if the response has no accountable owner, the organization has storage.
It does not yet have governance. Picture a normal week. Monday, a team creates a notebook for a vendor evaluation. They add a policy document, a proposal, meeting notes, and a spreadsheet of questions. Tuesday, someone changes the notebook visibility so another group can review it.
Wednesday, an administrator enables the log export to support reporting. Thursday, a privacy or security reviewer asks what the exported records contain. The team discovers actor emails, source names, source U R Ls, visibility states, network context, and artifact identifiers in a warehouse managed by a different group.
Friday, someone asks the ordinary questions. How long does that copy stay there? Who can query it? Does deleting the notebook delete the export? Can the log reveal the subject of the work? Who tells the users? And everybody looks at the architecture diagram, which contains exactly one cheerful box labeled A I.
Nothing dramatic happened. No breach. No villain. No robot monologue. The organization simply improved auditability faster than it updated the data map. Now replay the week with the receipt. The team states the purpose before creating the notebook.
It maps the content and control planes separately. It limits notebook visibility. It keeps the export off until an owner, retention rule, access group, and response use case exist. It runs one test event. It verifies what appears, who can see it, and what deletion does in both systems.
It gives users plain-language notice about the records created around the work. Same tool. Same audit capability. Very different operating maturity. Here is your forty-five-minute block for this week. Pick one A I workflow. Just one. A notebook.
An agent. A coding tool. A meeting assistant. A safety monitor. Anything that creates both work and evidence about the work. For the first seven minutes, write the approved purpose. What is the workflow for? Why is evidence needed?
Which secondary uses are not approved? From minute seven to fifteen, map the content plane. Sources. Inputs. Conversations. Outputs. Generated artifacts. Sharing. From minute fifteen to twenty-three, map the control plane. Actors. Events. I P data. Source metadata.
Labels. Alerts. Reviewer notes. Exports. From minute twenty-three to thirty-one, name every location and every role. Storage. Processing. Regional routing. Global storage. Warehouse. Administrator. Reviewer. Service account. Vendor. From minute thirty-one to thirty-eight, write the lifecycle. Creation. Lag.
Retention. Deletion. Hold. Backup. Schema change. From minute thirty-eight to forty-five, run one safe event through the path. Find the record. Verify access. Check the export. Name the response owner. Record the disposition. Then make one decision. Approved.
Approved with a control change. Limited pilot. Or held until the evidence path is governed. Do not try to map the entire A I estate before lunch. That is how forty-five-minute blocks become three-quarter road maps and quietly die in a steering committee.
One workflow. Two planes. Seven receipts. One decision. Auditability and privacy are not opponents. A well-designed system can produce evidence without making that evidence available to everyone forever. It can support safety monitoring without turning every administrator into a content reviewer.
It can help investigate misuse without converting a reporting warehouse into an unbounded record of worker activity. It can separate content from signals, and still preserve a responsible path for review and response. But that outcome does not come from the word audit.
It comes from architecture, purpose, access, lifecycle, and tested ownership. The log is not automatically the safe part of the system. The export is not automatically the neutral part. The receipt is not automatically harmless because it proves control.
If you cannot govern the receipt, the receipt is part of the risk. That is the check for this week. Map the content. Map the control evidence. Then prove that the people, places, permissions, and timelines match the purpose you approved.
Because the prompt is data. The output is data. And the audit trail is data too. This has been AI Change Desk. I am Michael. Thank you for listening. Take care of the people doing the work, and govern the records created around them.
I will see you next week.