Someone asks, “Find me a pair of shoes I can walk in all day without looking clunky.” Once, tossing this plea into a search bar yielded ten pages of blue links and a half-page of sponsored ads. Today, whispering it into a chat window typically returns a single name, a singular rationale, and a quiet verdict: “I recommend this one.”
The moment that name is minted, the underlying logic of the marketplace quietly shifts its tectonic plates. We used to wage war for clicks; now, the battlefield is trust.
The philosopher Baruch Spinoza famously posited, “Omnis determinatio est negatio”—to define is to exclude. Once a concept is definitively named, the reality of all other options is abruptly severed. AI shopping agents are breathing life into this philosophical maxim every single day. They christen “the one worth buying” on our behalf, neatly omitting why the other ninety-nine are deemed unworthy. Consequently, merchants are waking up to a glaring void in their analytics dashboards. The question is no longer, “Why isn't anyone finding me?” but rather the far more haunting, “Why wasn't I recommended?”
Over the past year or two, the industry has fractured this phenomenon into a handful of buzzwords like AEO and GEO. They speak of grooming product data, maintaining entity consistency, and securing third-party citations so that AI can more easily “read, cite, and crown” a product. Industry think pieces tirelessly iterate: titles styled as catchy slogans pale in comparison to extractable facts; conflicting brand narratives scattered across the web severely erode a model's confidence; and off-site reviews or editorial accolades often bear more weight as “evidence” than a meticulously crafted landing page. All of this is true, yet it largely reads like an extension of the old SEO playbook—dress your content up nicely, and wait to be noticed.
However, running parallel to this quest for “AI visibility” is a far more profound endeavor: empowering AI to actually execute the purchase for the user and manage operations for the merchant. Anthropic recently unveiled a blueprint and showcase for commerce agents. On one end, a shopper agent scouts and compares; on the other, a merchant agent minds the store and updates catalogs. They even provided a litany of common skills, safety guardrails, and blueprints for tethering large language models directly to a store's proprietary inventory and order management systems. The revelation here isn't the sheer brilliance of a singular product, but rather that the entire architecture of “how to build an agent from end to end” has been commoditized into practically plagiarizable templates. Shopify is echoing this exact trend, reporting a massive surge in AI-driven traffic and utilizing MCP (a universal protocol connecting models to external systems) to read real-time inventory and place actual orders. Gartner has even prophesied that by the close of 2026, a substantial swathe of enterprise software will harbor these “single-task, purpose-built agents.”
As the mechanics of running an agent become commoditized, the true scarcity will no longer lie in the quick-witted banter of a chat interface. The true rarity resides in the bedrock beneath it: what deserves to be trusted, what merits comparison, what justifies a recommendation, and what actions are ultimately permitted to take place.
Though this sounds abstract, it is intensely practical. A truly responsible recommendation must survive a gauntlet of trials. Was the core desire accurately understood (separating hard constraints from mere preferences)? Have the candidates been rigorously filtered by the harsh realities of commerce (is it in stock, can it be shipped, does it break the bank)? Are the claims bolstered by an anchor of evidence (are the sources independent, current, and harmonious)? And finally, can the rationale be replayed—not as a beautifully spun post-hoc rationalization, but as an auditable trail of breadcrumbs left at the exact moment of decision?
Historically, merchants have consoled themselves with “visibility scores,” much like taking one's temperature. A thermometer is useful, but a fever is a symptom, not the disease. The genuinely arduous problems are structural, resembling a funnel of failure: Was I never recalled in the first place? Was I recalled but eliminated due to low inventory? Did a lack of evidence bar me from the comparison stage? Or did I simply lose the final duel to a more verifiable competitor? Conflating these distinct failures into a nebulous “insufficient AI score” ensures that optimization efforts will forever strike the wrong target.
As Machiavelli might have cautioned regarding statecraft, without a measurable metric for trust, any system of reward and accountability is utterly baseless. Translated into the vernacular of agentic commerce, this metric takes the form of explainable decisions and auditable executions. A model may propose, but it must never assume the mantle of ultimate authority; the sovereign truth of pricing and inventory must firmly remain within the merchant's own citadel. In their public documentation, Anthropic deliberately underscores that Claude serves as the intelligence layer, not the storefront itself—relationships, catalogs, and fulfillment remain the undisputed domain of the merchant. This is a subtle yet profound wake-up call to the entire industry: the digital shell may be open-sourced, but the facts, evidence, and control mechanisms beneath that shell are the true battlegrounds of long-term competition.
Therefore, if we treat “being seen by AI” as the ultimate destination, we merely strand ourselves halfway up the mountain of optimization. The more holistic question is this: In a world where agents act as proxies for discovery and comparison, how does a merchant build an intelligence, control, and observability layer that can be reused by multiple, disparate agents? A layer that ensures products are accurately understood, credibly compared, and granted a quantifiable chance at being recommended. Outwardly, this could be branded as a “Merchant OS for AI Product Discovery.” Inwardly, however, it is simply a more humble promise: a recommendation is not black magic; it is a judgment that can be questioned, scrutinized, and continuously refined.
The philosopher Ludwig Wittgenstein famously argued that once a ladder has been used to climb to a higher plane, it must be thrown away. The technological tools of our trade will inevitably become cheaper, eventually fading into the background like the air we breathe. What will be remembered is the destination we reached—why a recommendation was forged, why a rejection occurred, and whether, the next time around, we can make the process just a little bit fairer.
