A shopper asks ChatGPT for a coffee maker under two hundred dollars. Not on Google. Not on Amazon. Not on your site. Inside ChatGPT. And that matters because the first serious comparison now happens before the shopper ever reaches your storefront. This is AI Change Desk, episode fourteen: Commerce Surface Check. If EP011 was about AI disappearing into the work surface... and EP012 was about whether anyone can actually see what that change is doing... this is the customer-facing version of the same shift. A new surface matters most when it changes where decisions begin. And that is what happened this week. On March twenty-fourth, OpenAI said it improved shopping in ChatGPT so users can discover, compare, and decide on products more easily. The company says results are more visually rich, easier to compare side by side, and better on data coverage, freshness, and speed. OpenAI also said these improvements are powered by an expansion of the Agentic Commerce Protocol for product discovery. That is the first signal. The second signal is the merchant side. Also on March twenty-fourth, Shopify said millions of merchants can now sell to ChatGPT users through what it calls Agentic Storefronts. And the operator detail there is not just distribution. It is control. Shopify says orders can carry ChatGPT referral attribution... and that merchants remain the merchant of record. That is a much more operational statement than, “AI shopping is growing.” And then there is the concrete retailer example. Sephora announced an app in ChatGPT that helps customers discover products with recommendations tied to their Beauty Insider profile if they choose to link it. Sephora also said users can use loyalty rewards and certain benefits now... with payments and checkout inside the app planned for a future update. So put those together and the pattern gets pretty clear. Discovery is moving. Not all commerce. Not all checkout. Not all intent. But discovery, comparison, and early narrowing of options are starting to happen in a conversational AI surface before the shopper reaches the place most teams still treat as the real beginning of the funnel. And that is where operators need to wake up a little. Because if discovery moves upstream of your site, then your old measurement habits get weak very quickly. You can still measure site traffic. You can still measure cart conversion. You can still measure paid performance. But if a customer already got half the decision made inside ChatGPT, then your site analytics are only seeing the second half of the story. That is the value here. The value is not that OpenAI added prettier shopping cards. The value is that product discovery is becoming an AI-native operating surface. And when a new surface becomes real, the teams that win are usually not the ones with the most excited launch post. They are the teams that figure out what actually needs to be instrumented. So what should operators care about this week? Three things. First: product data quality just got more consequential. If a conversational system is deciding what to surface, compare, summarize, and rank, then clean titles, attributes, images, availability, and price freshness are no longer a catalog hygiene issue sitting off in the corner. They are discoverability infrastructure. That is one reason the Shopify language matters. It is not describing AI chat as a novelty plugin. It is describing AI channels as another commerce surface that needs centralized product and order infrastructure behind it. Second: attribution is about to get messier... and more important. If a platform can tell you a shopper came from ChatGPT, good. That is useful. But it is not enough. You still need to know whether the product was surfaced because your feed was clean, because your merchant ranking was stronger, because the shopper asked a highly specific question, or because the category is simply easy for the model to compare. So “AI referral traffic” is not a full answer. It is the beginning of the answer. Third: merchandising and channel teams are going to collide here. Because this is where teams usually get it wrong. They treat the whole story like PR or SEO. They say, “great, we need to show up in ChatGPT.” Okay. Fine. But that is only the outer layer. The deeper question is whether your internal systems are good enough for AI-native discovery. Do you have clean attributes? Do you have current price and availability? Do you have a clear path from product detail to merchant detail to conversion? Do you know which categories are safe for this surface... and which ones are going to create support headaches? That is the chapter where teams usually get uncomfortable. Because “be visible in AI” sounds strategic. “Fix the product feed, align attribution, and decide which categories are ready” sounds operational. And operational... is where the real work is. Here is the mini-case I want people to keep in their head. Take something simple, like small kitchen appliances. A shopper says, “I want an espresso machine under four hundred dollars that is easy to clean and does not take up too much counter space.” That is already a structured buying question. If ChatGPT can compare options side by side and narrow the field before the shopper clicks out, then your advantage is not just brand recognition. Your advantage is whether your product data helps you survive the comparison. If your dimensions are inconsistent... if your cleaning features are vague... if your images are weak... or if your pricing is stale... you are not losing at checkout. You are losing before the shopper ever opens your page. That is a different failure point. And that is why this week matters. Now, the hidden tradeoff. There is a temptation to hear all of this and think, “great, a higher-intent shopper arrives pre-qualified.” Maybe. But the tradeoff is concentration. The more discovery happens upstream in a small number of AI surfaces, the more pressure shifts to whoever controls the comparison layer, the ranking logic, the metadata quality, and the merchant-selection rules. OpenAI’s own shopping help documentation is useful here. It says product results are not ads and are not influenced by partnerships. It also says merchant rankings can depend on things like availability, price, quality, whether the merchant is the maker or primary seller, and whether Instant Checkout is enabled. That is helpful. But it also means operators cannot treat this like a simple paid-placement model they already understand. You may get more intent. You may also get less clarity about why some products surface better than others unless your data, category strategy, and merchant signals are disciplined. So what would I decide by Friday? I would not start with the whole catalog. I would pick one category where comparison behavior is already common. Appliances. Beauty. Consumer electronics. Home tools. Something like that. Then I would do five things. One: pick the category and define the actual buying question you care about. Not “AI commerce.” A real query. “Best noise-canceling headphones for travel under three hundred dollars.” Two: audit the product data that a conversational system is most likely to lean on. Title quality. Attribute completeness. Pricing accuracy. Review coverage. Availability. Images. Three: define an attribution view specifically for AI-originated demand. Not because it will be perfect. Because if you do not define it early, every internal conversation becomes vibes. Four: decide whether you are only optimizing for discovery... or whether you are willing to support deeper merchant-side experiences. The Sephora example matters because it shows the difference. Discovery is one level. Linked loyalty, benefits, and future in-app checkout are another. Five: pick one owner. If merchandising thinks this is ecommerce’s job, and ecommerce thinks it is growth’s job, and growth thinks it is search’s job... nothing useful happens. This needs one operator who can force a cross-functional answer. And one thing I would not do. I would not measure success by mentions alone. Not by “we show up in ChatGPT.” Not by screenshots. Not by one nice-looking recommendation result. That is the early-stage version of dashboard theater. The question is not whether you can be seen. It is whether the new surface changes qualified discovery, category conversion, or merchant economics in a way you can actually defend. That is the threshold. So here is the question for this week. If product discovery starts in AI before it starts on your site... does your team know what to measure first?