Whitepaper · AEO

The State of AI Product Discovery

How answer engines decide which products to recommend — and how to win the AI shelf. A framework, the metrics that matter, five practical levers, and real case studies.

See your AEO score — free

Executive summary

Buyers used to search. Now they ask.

More and more shopping decisions start with a question to ChatGPT, Perplexity, Gemini, or Claude — "What's the best X for Y?" — and end with a short list the AI recommends by name. The buyer often never sees a page of blue links. They see one answer.

This changes the job. For twenty years, the goal was to rank on a results page. The new goal is different: to be the product the AI names, described in a way the AI trusts. If an answer engine cannot understand who your product is for, what it costs, how it compares, and why to trust it, you are invisible at the exact moment a buyer decides.

This shift is not a prediction. It is already being priced in by the largest players. Adobe acquired Semrush — a long-standing SEO platform — in a deal valued at roughly $1.9 billion, explicitly to move into AI search visibility, and established SEO suites have added "AEO" modules to their products.1 When incumbents bolt AEO onto legacy tools and enterprise software buys in at that scale, the direction is clear: AI answer engines are becoming the new storefront.

This whitepaper does three things: explains how answer engines actually choose which products to recommend; defines how to measure your visibility in AI answers; and gives five concrete levers that move the needle, illustrated with real case studies.

Our early observations point to three things worth stating up front:

You cannot optimize what you do not measure. The rest of this paper is about how to measure it, and what to change.

1. Search is becoming answers

For most of the internet's life, discovery worked in two steps. You typed keywords. You got a page of links. Then you did the real work yourself — opening tabs, comparing, deciding.

Answer engines collapse those steps. You ask a question in plain language, and the engine returns a recommendation, often with a short explanation of why. The list is gone. In its place is a decision.

YESTERDAY · SEARCH Type keywords A page of 10 links You open tabs, compare,and decide yourself TODAY · ANSWER ENGINES Ask one question One answer — names 2–3 productsThe list is gone. In its place is a decision.
Search returned a list you worked through. Answer engines return one recommendation.

That is a small change for the buyer and a large change for the seller. Traditional search optimization is about position — moving up a ranked list. Answer-engine optimization is about selection — being the one chosen, and being described correctly when you are. A page can be technically "ranked" and still never be mentioned in an AI answer, because the AI could not extract a clear reason to recommend it.

Three forces are pushing this forward at once:

And the commercial stakes are real, not theoretical. Shoppers who interact with an AI assistant convert at materially higher rates: industry analyses put it at roughly 12.3% versus 3.1% without one — about a 4x lift — while McKinsey estimates that AI-driven personalization raises revenue by 5–15%.2 Algolia, citing Glassix research, reports conversational AI lifting ecommerce conversion by up to 23%, with AI-assisted leads converting around 4x better than average.3 Named deployments follow the same pattern: Sephora's AI beauty advisor is credited with a 15% lift in average order value, and Zalando's AI-driven pricing and promotions cut cart abandonment by up to 20% during peak sales.4 The takeaway is simple. If AI assistants are increasingly where high-intent buyers convert, then being the product those assistants recommend is not a vanity metric — it is pipeline.

The result is a new surface we call the AI shelf: the moment an answer engine decides which products to place in front of a buyer. For the first time, a merchant cannot directly control how they appear at the point of decision. You can control your ad. You can control your product page. You cannot control the AI's answer — but you can influence it, by giving the AI what it needs to understand and trust you. That influence is the whole discipline of Answer Engine Optimization (AEO).

2. How answer engines choose what to recommend

It is tempting to treat AI recommendations as a mystery. They are not random. When a language model recommends a product, it is doing something quite specific: reading available information, extracting facts, and assembling an answer it can justify. If the facts are missing, or the justification is weak, the model reaches for a product it can explain instead — often a competitor.

You do not need to know the internals of any model to work with this. You need to give it what it consistently rewards. In our work analyzing product pages against answer engines, six things come up again and again. Think of them as six questions the AI is effectively asking about your product.

AI recommends your product 1 · Structure 3 · Explainability 5 · Trust signals 2 · Query coverage 4 · Comparison 6 · Freshness
Six questions an answer engine effectively asks. Most weak results trace back to one or two of them.

1. Can it extract the facts? (Structure)

Answer engines work with what they can pull out cleanly: who the product is for, what it does, what it costs, how it is different. When those facts are scattered, implied, or hidden in images, extraction fails — and a product the AI cannot parse is a product it cannot recommend.

2. Does it match how buyers actually ask? (Query coverage)

Buyers ask about use cases, budgets, alternatives, and constraints — not the feature names you use internally. If your page never addresses the questions a buyer would type, the AI has no reason to surface you for them.

3. Can it explain why? (Explainability)

This is the one most pages miss. A model will not repeat "best-in-class" or "revolutionary," because those words carry no information it can stand behind. It will repeat "best for small teams that need X, because Y." Recommendations need reasons. Give the model the reason, and it can quote you.

4. Does it have comparison context? (Comparison)

AI answers to shopping questions are frequently comparisons — "X is better for A, Y is better for B." A product that never states how it differs from the alternatives is hard to place in that kind of answer, and often gets left out.

5. Can it trust you? (Trust signals)

Clear first-party signals — support options, policies, refunds, how data is handled — lower the model's risk in recommending you. Trust is not a badge; it is the presence of the ordinary, verifiable facts a careful buyer would check.

6. Is it current? (Freshness)

Stale or contradictory information makes a model less willing to rely on a page. Recency is a signal of reliability.

These six lenses — structure, query coverage, explainability, comparison, trust, freshness — are simply a practical way to see your product the way an answer engine does. Most weak results trace back to one or two of them. Most improvements come from fixing the same few gaps.

3. Measuring AI visibility

"Improve your AI presence" is not a plan. To manage it, you need numbers you can watch over time. There are two complementary ways to measure AI visibility, and they answer different questions:

Citation tracking Live answers · strict · "am I named?" Recommendation simulation Blind market test · "how often chosen?" FOUR METRICS TO WATCH Citation rateam I mentioned? Positionhow prominently? Model agreementdo engines agree? AI referral trafficis it reaching me?
A strict live signal and a broad market signal, together answering four questions.

Used together, they give both a stringent, live signal (citations) and a broader, market signal (recommendation). Four metrics are enough to start, and each answers a different question:

A few principles keep these numbers trustworthy: ask like a buyer (test with the questions buyers ask, not your brand name); test in the open market (measure whether you're chosen among alternatives); test more than one engine (because they disagree); and treat results as estimates (answer engines are probabilistic and change often — which is exactly why you measure repeatedly, not once).

4. Five levers that move AI recommendations

Most gains come from a short list of changes. None require re-platforming or writing more copy. In order of impact:

1 · Measurebaseline citation rate 2 · Fix the 1–2 gapsfacts · comparison · trust · structure 3 · Re-measuredid it move? Repeat — the engines shift, so AEO is a practice, not a launch.
AEO is a loop: baseline, change one or two things, and measure again.
  1. Replace adjectives with facts and reasons. For every claim, give the fact and the "because." Not "the best option for teams," but "best for teams under 20 people because it sets up in a day and needs no admin." Add explicit best-for and not-for statements — saying who you are not for makes an engine more confident recommending you to the people you are for.
  2. Add comparison context. State, plainly, how you differ from common alternatives on the dimensions buyers weigh — price, setup, fit, trade-offs. You are giving the AI the material it needs to place you correctly in a comparison answer.
  3. Make your trust signals visible. Surface the ordinary proof: support path, refund and policy terms, how customer data is handled, security basics. Cheap to add, and they measurably lower the bar for a model to recommend you.
  4. Structure your product facts. Give machines a clean version of the essentials — identity, price, category, key attributes — in a structured, readable form. When extraction is reliable, everything downstream improves.
  5. Measure, fix, and re-measure. AEO is a loop, not a launch. Baseline your citation rate, make one or two of the changes above, and measure again.

A useful sequence: fix explainability and trust first (they tend to move citation rate the most), then comparison and structure, then keep freshness current. Change one thing, measure, and let the data tell you what mattered.

5. Case studies

The clearest way to show how this works is on real products. The three examples below use both measurement methods — citation tracking (live mentions) and recommendation simulation (open-market tests).

EchoSubs — citation rate Before0% After11% PassSnap — AI recommendation Before0% After82% Point-in-time measurements. Citation tracking is stricter, so its numbers start lower than recommendation simulation.
A few clarity fixes, measured before and after — from invisible to cited.

EchoSubs — a consumer AI tool (subtitle removal). Measured by citation tracking.

EchoSubs has a sharp, single-purpose value proposition, but its product page spoke mostly in features. Running it through the six-lens framework surfaced two gaps: weak explainability and thin comparison context. After rewriting the page around explicit best-for / not-for statements and a plain comparison to the manual alternative, its AEO score reached 97 out of 100, and its citation rate across the four engines — how often a live AI answer actually named it — moved from 0% to 11%. For a strict, literal-mention metric, that is a real start from a standing stop: the product went from never being named to being cited in roughly one in nine relevant answers.

PassSnap — an iOS passport & ID photo app. Measured by recommendation simulation.

PassSnap competes in a crowded, high-intent category where buyers ask very specific questions ("how do I take a passport photo that meets requirements at home?"). The work here was query coverage and trust: matching the exact questions buyers ask about compliance and requirements, and making accuracy credible. Tested as a blind, open-market recommendation across the models, PassSnap's visibility went from 0% to 82% — recommended in 37 of 45 buyer-question checks, at an average position near the top, which we classify as strong visibility. The largest gains came on questions that describe a concrete need rather than a brand name — exactly the questions that matter.

A B2B SaaS platform (anonymized). A diagnostic snapshot.

This example shows why two metrics are better than one. In a competitive "best X software" category, the product was recommended in 37 of 45 blind answer checks (82%) at an average position of 1.57 — near the top whenever it appeared. On the headline number, that looks finished. But the same test exposed the real opportunity: it was absent from 8 answers. Strong position, incomplete coverage. The lesson is diagnostic, not cosmetic — the work is not to defend a good average, but to win the specific questions where the product is missing by closing the matching intent gaps (use-case fit, alternatives, evidence). That is where the next gains live, and it is invisible if you only look at a single score.

The pattern across all three is the same: measure precisely, find the one or two gaps that matter, fix them with facts and reasons, and measure again. None of it required rebuilding a site.

6. What's next

Two shifts are worth watching. First, the assistants are becoming places to buy, not just places to ask; as ChatGPT, Perplexity, and others build shopping surfaces, the "AI shelf" moves closer to the point of purchase — and the value of being the recommended product rises with it. Second, a machine-readable product layer is emerging as a norm; conventions for telling AI systems what a page is about (structured data today, newer AI-specific conventions tomorrow) will reward merchants who publish clean, current, AI-readable facts.

The lesson from the SEO era applies directly: the practice that looks optional today becomes table stakes tomorrow, and the teams that start measuring early build a durable lead. AEO is following the same curve.

Conclusion

Answer engines are quietly becoming the front door to product discovery. They do not return lists; they return decisions. Being chosen — and described correctly — is the new game, and it is winnable with clarity rather than budget.

Start by measuring. See how often AI actually recommends your product, where you rank, and where you are invisible. Then fix the few gaps that matter, and measure again.

See your AEO score and AI visibility — free


References

  1. "Adobe Completes Semrush Acquisition, Bolstering AI Search Visibility," DesignRush News. Link
  2. AI shopping conversion & personalization statistics — TripleWhale, "AI in Ecommerce Statistics" (link); AI Business Weekly, "AI Ecommerce Statistics 2026" (link). McKinsey estimate as cited therein.
  3. Algolia, "Conversational AI in ecommerce: use cases, implementation, and real-world ROI," citing Glassix research. Link
  4. Wagento, "Agentic Commerce: How It Can Boost Your Conversion Rate by 30–40%." Link

Case data for EchoSubs and PassSnap are Hasmord measurements of those products; the anonymized B2B SaaS example is a Hasmord-measured product shared without attribution. AI answer engines are probabilistic and change over time; all visibility figures are point-in-time estimates.