On September 16, researchers from Maastricht University, Utrecht University and the University of Zurich posted a new paper on arXiv. Its title sounds exactly like a real shopper: “If I Had to Buy Just ONE: Galaxy S26 Ultra”.

They didn’t invent questions out of thin air. They started from 2,528 real requests for buying advice and narrowed them to 117 prompts asking for a physical product. They sent each prompt to ChatGPT and Gemini, through both the logged-out chat interface and the API, and to Google’s AI Overviews, three times, hours apart. The goal was not to find out which assistant was smarter. They wanted to know whether the recommendations were stable, how confident the assistants sounded, and whether the sources they cited lined up.

Same Prompt. Different Picks. Why bot recommendations change. A shopper asks “Best wireless headphones under $200?” and three AI assistants, labelled AI Alpha (high confidence), AI Beta (moderate) and AI Gamma (lower confidence), each recommend a different pair of headphones, citing different sources with a different one-line verdict.
One prompt, three assistants, three different picks. Illustration.

The numbers speak for themselves. In answers that recommended a product, ChatGPT expressed a first-person preference, along the lines of “my pick would be…”, 79% of the time. Gemini did so 7% of the time, and AI Overviews 2%. For the same prompt, the websites cited by the ChatGPT and Gemini interfaces overlapped by only 5.4% on average, and in 76.7% of comparisons they shared no website at all. Ask the same question again and the recommended products often change. The public interfaces and the developer APIs don’t match either. Even when the same prompt went to ChatGPT’s interface and its API within a fraction of a second, the two shared only 12% of their cited domains on average (14.8% for Gemini). The paper’s conclusion is blunt: neither a single response nor API data can be assumed to represent the advice consumers actually see.

There’s a line often attributed to Voltaire: if you want to talk with me, first define your terms. In AI shopping, the terms are more than a model number. They include the assistant’s tone: “this is my pick” versus “you might consider this.” They include the links cited next to the answer, and whether the same product still shows up when you ask again a few hours later. When those terms keep shifting, the recommendation blurs. The buyer thinks they got a firm answer. What they may have gotten is one roll of the dice.

What does this mean for retailers? Stop arguing about which assistant is biased; that’s tech gossip. The harder truth is simpler. When a shopping choice happens inside a chat window, your product isn’t sitting on a fixed shelf waiting to be picked. Different systems, with different tones and different sources, keep reinterpreting it and comparing it with others. A screenshot of your product being recommended proves it happened once. It doesn’t prove it will happen again, and it certainly doesn’t prove the assistant described your specs and limits correctly when it made the pitch.

Real buyers ask questions far more specific than “what’s the best one?” Can I use this in the rain? Does the test data match this exact model year? What does the warranty actually cover? Who is this product a bad fit for? If it’s two centimeters wider, will it still fit in my cabinet? Shoppers used to ask these in customer-service emails or explain them on return forms. Now assistants check these details up front. A detail your product page leaves blank gets magnified. A page packed with text that leans on outdated marketing copy and old test results just feeds the assistant bad data. When a buyer asks and your data can’t answer, that gap is a missed buying opportunity. It isn’t one lost sale. It is a whole category of demand you never addressed.

The industry likes to treat answer engine optimization as a finish line: get cited, and you’re done. Those tactics are still useful. But this study shows the flaw: citations and recommendations are unstable by nature. What works is a routine. Find the missing information. Add hard facts that match the exact item. State the limits clearly. Link every strong claim to evidence for the current model. After you fix a page, ask again, in the public chat apps too, and see how the product is understood and compared now. Optimization is like washing windows: the glass always gets dirty again. People phrase questions differently, the tone shifts, the cited sources change. Make observation, evidence and clear writing a habit, and stop treating a lucky screenshot like a quarterly report.

Think of it this way. Comparison shopping used to mean looking at the same catalog page side by side. Typing “pick one for me” today is like walking into three different stores. One clerk hands you a top pick. Another stops to ask about your budget and where you’ll use it. A third points you to a different set of reviews. Different storefronts are normal. What’s strange is that merchants update one web page and never check how these assistants actually talk about their products: whether the item was misdescribed, fairly beaten, or ignored.

When you hear new questions from real customers, feed the answers back into your product information. That’s a principle, not a shortcut. What works is a list of verifiable facts and the discipline to check again after you hit publish.

That is the work Hasmord, the AEO-native sales platform, is built for. It doesn’t fight over checkout protocols, and it doesn’t promise ranking boosts. It helps brands find missed buying opportunities: the real buyer questions their product information can’t answer yet. It checks whether the evidence behind each claim matches the specific item. And it sends buyer questions to AI models through their APIs, keeps every answer with the model and the date, and repeats the checks over time, so a team sees a pattern rather than a single screenshot. As the paper shows, an API answer is not what the consumer app shows, so we label it for what it is; it is worth pairing with a look at the public chat apps. The goal is for the machine to understand the product, and for the buyer to have a factual reason to choose it. Back to the paper, the point is simple. The same prompt already produces different answers. Can your product data survive being asked the same question twice?

Typing “pick one for me” isn’t the end of the shopping journey. It shifts the pressure back to the digital shelf. If assistants give unstable answers, buyers will push even harder on vague product descriptions. The quiet teams will sit down with their data gaps, their evidence and the way they check results. The loud ones will keep posting screenshots of a lucky mention. A screenshot expires. What buyers remember is whether you had a verifiable fact ready when the assistant compared you with the competition.