How to reverse-engineer where AI recommends the other guy, and where that stops being useful.

Don't start by inventing thirty prompts you think favor your product. Start by finding the exact buyer conditions where the recommendation flips. That boundary is the challenger's actual positioning, discovered rather than brainstormed.

Don't guess the boundary.
measure where it moves.

A founder building a challenger product asked me to help him understand something specific: under what circumstances does ChatGPT, Claude, Gemini, or Perplexity recommend his product over the much larger incumbent he competes with.

His first instinct, and the instinct I see from almost every founder who asks this question, was to write thirty AEO prompts that he expected his product to win, run them, and count the mentions.

I told him I would not start there. Prompts you invent yourself encode the answer you already believe. The interesting question is not "can I get my product mentioned." It's this one:

The actual research question

Under what buyer circumstances does an LLM naturally conclude that the challenger is the better answer than the incumbent? Those circumstances are the challenger's defensible positioning, discovered empirically rather than brainstormed.

This post is the method I used, written up as a tactical guide. It generalizes past that one project: any time you want to know where AI is quietly recommending a competitor over you, or where you should stop competing entirely, this is the design.

Call it a recommendation boundary test

The idea is simple. Picture a line between two products:

Incumbent ← buyer context → challenger

Somewhere along that line, for some combination of buyer needs, team size, workflow stage, and persona, the recommendation flips. Your job is to find where, across real models, not to assert where you think it should be.

You do this by starting broad and unbranded, then adding one constraint at a time until the answer changes.

Step 1: Start with jobs, not brand names

Don't mention either product yet. Ask the job-to-be-done the way a buyer actually would.

For a design-prototyping category, that looks like:

This establishes unaided discovery. For every prompt, on every engine, log four things: does it mention the incumbent, the challenger, both, or neither, and in what order.

Resist the urge to skip this step because you already "know" who wins here. You don't. You're building the baseline the rest of the test measures against.

Step 2: Add the attributes the challenger claims differentiate it

Now manipulate the buyer's requirements one variable at a time. This is where the useful signal starts showing up.

Start with the plain job:

"What's the best AI tool for prototyping a SaaS feature?"

Then layer constraints, one prompt per addition:

Then combine all of them into one buyer-realistic scenario:

"We're a 50-person SaaS company. We have an existing design system and want PMs and designers to prototype new features together, collect stakeholder feedback, preserve previous iterations, and then hand the approved design to engineering as production code. What tools should we evaluate?"The compound prompt — this is the one that matters

This tests whether the challenger's claimed differentiation actually shows up in recommendation behavior, or whether it's marketing copy that hasn't earned a place in the model's reasoning yet.

Step 3: Deliberately test where the incumbent should win

This step gets skipped constantly, and it's the one that keeps the research honest.

Test the conditions favoring the incumbent just as hard:

If the incumbent consistently wins these, that's not a gap to close. That's a boundary to respect. You're building a recommendation map, and the LLM results fill it in, not your assumptions.

Buyer needIncumbentChallenger
Quick one-off prototypeStrongStrong
Already inside the incumbent's ecosystemStrongModerate
Solo founderStrongModerate
Existing design system??
PM + designer collaboration??
Stakeholder comments??
Extensive version history??
Production handoff??
Long-running product system??

Leave the question marks as question marks until you've run the prompts. Filling them in yourself defeats the entire point of the exercise.

Step 4: Test buyer maturity, not just features

The most interesting positioning I've found running this test was never a feature. It was a stage of work.

Hypothesis: the challenger becomes more appropriate as the prototype becomes more consequential. Test it as a staged sequence:

StagePrompt
1. Idea"I had an idea for an app this morning and want to see what it could look like. What should I use?"
2. Prototype"I'm a PM and need to show my team three approaches to a new feature. What should I use?"
3. Collaboration"Three PMs and a designer need to iterate on an AI-generated prototype together. What should we use?"
4. Stakeholder review"We need executives and customers to review and comment on several prototype directions."
5. Engineering handoff"We've approved the prototype and need to hand it to engineering without rebuilding everything from scratch."
6. Ongoing system"Our product organization wants AI prototyping to become part of how we develop features every week across multiple teams."

Watch where the recommendation changes across the six stages, not whether it changes once. If the incumbent dominates stages 1 and 2 while the challenger increasingly appears from stage 3 onward, you've found something worth building a positioning statement around, something like: the incumbent helps you make something; the challenger helps a team develop a product. That's a claim you can now defend with evidence instead of one you invented in a positioning workshop.

Incumbent share of recommendationChallenger share of recommendation
1 · Idea
2 · Prototype
3 · Collaboration
4 · Stakeholder review
5 · Engineering handoff
6 · Ongoing system
Stage 1Stage 6

Illustrative shape of the pattern this test looks for. Your actual split comes from the runs, not this chart.

Here's what that looks like in a real run. Same engine, same logged-out conditions, two stages of the same product journey:

Gemini answering the Stage 1 prompt 'I had an idea for an app this morning and want to see what it could look like': it recommends Figma and Balsamiq for mockups, Bolt.new, Lovable and Magic Patterns for AI-generated apps, and Whimsical or FigJam for brainstorming Gemini answering the Stage 5 prompt about handing an approved prototype to engineering: it recommends Figma Dev Mode and Zeplin, Builder.io and Anima, then Storybook and shadcn/ui
Stage 1 (idea) vs. Stage 5 (engineering handoff). Figma is the only name in both answers, and in Stage 5 it shows up as Dev Mode. Everything else on the shortlist changes: Bolt, Lovable and Balsamiq drop out, and Zeplin, Builder.io and Anima come in. Gemini, logged out (Flash-Lite), September 24, 2026. One run on one engine, so read it as the shape of the signal, not a result.

Step 5: Add two more dimensions — company size and persona

Run the same job prompt while changing exactly one thing each time.

Company size: solo founder, 5-person startup, 20-person startup, 50-person SaaS company, 200-person SaaS company, 1,000-person enterprise. Ask the same compound prompt from Step 2 with only the headcount changed. Where does the challenger begin winning? That answer is close to the challenger's real ICP, whether or not it matches who the team believes they're selling to.

Persona: founder, product manager, product designer, engineering manager, head of product, CTO. Keep the scenario identical and change only who's asking. A challenger can have wildly different recommendation rates by role, and that's commercially useful information a mention-rate dashboard will never surface.

Step 6: Ask the model why, every time

This is the step that turns a mentions count into positioning research. After every recommendation, append:

What specifically made you recommend [X] over [Y]? Cite the evidence supporting each reason. Separate product facts from your own inference.

Now you're not measuring mentions. You're extracting the reasoning attributes tied to each recommendation. After enough runs, patterns emerge that read like a positioning brief: the challenger wins when the model detects collaboration, an existing design system, PM ownership of the workflow, and an engineering handoff downstream. The incumbent wins when it detects an individual, speed, a one-off task, and someone already inside that ecosystem. That's positioning you can defend, because it came from the reasoning, not from a workshop.

Gemini's answer to the why follow-up, split into Product Facts with cited sources for Figma Dev Mode and Builder.io, and Inferences and Reasoning covering workflow continuity and the maintenance risk of generated code
The "why" follow-up after the Stage 5 answer. The model separates cited product facts from its own inferences, and the inferences are the useful part: "keeps the engineering team inside their preferred workflow" and "maintenance risk of automated code" are reasoning attributes you can count across runs. Gemini, logged out (Flash-Lite), September 24, 2026.

Step 7: Run it across models, and run the important ones twice

Use the same prompts against ChatGPT, Claude, Gemini, and Perplexity, in clean, unauthenticated conversations so history doesn't bias the answer. Recommendation output varies run to run, so for the compound and staged prompts that matter most, run each one more than once before you trust the result.

Build a simple matrix: prompt, persona, company size, requirement, and the pick from each engine. The goal isn't a spreadsheet with a hundred rows. It's a small number of well-chosen rows you can actually reason about.

Step 8: Run it twice more, with and without the vendors' own words

This is the modification I'd make to almost every version of this test I've seen done before, including earlier versions I ran myself.

Run two conditions on every important prompt.

Condition A, natural environment: "Recommend the best tools for..." with no restriction. This measures what an actual buyer sees.

Condition B, independent evidence only: the same prompt, plus "Do not use information published by either vendor as evidence. Base the recommendation on independent third-party sources."

Comparing A and B is far more diagnostic than a raw mention count:

PatternWhat it means
Wins A, disappears in BThird-party evidence problem — the recommendation is riding on the vendor's own marketing, not independent proof
Loses A despite strong evidence in BDiscoverability problem — the proof exists but isn't reaching the model's retrieval
Wins bothDefensible position — amplify it
Loses bothPositioning or product evidence problem — the claim isn't supported yet

The difference shows up even in a single run. Here's the Stage 5 handoff prompt from earlier, run again with the independent-evidence restriction added:

Condition A: Gemini's unrestricted answer to the handoff prompt, naming Figma Dev Mode, Zeplin, Builder.io, Anima, Storybook and shadcn/ui Condition B: the same prompt with 'do not use information published by any tool's vendor as evidence' added; Gemini answers with practices like design tokens and living specifications, cites UXPin, BrandyHQ and Qt, and names only Figma's Dev Mode as a product
Condition A (left) vs. Condition B (right). Once vendor-published evidence is excluded, Zeplin, Builder.io, Anima and Storybook disappear. The answer turns into process advice, and the only product left is Figma's Dev Mode, now backed by a third-party source. That's the "wins A, disappears in B" pattern: those recommendations were riding on the vendors' own pages. Gemini, logged out (Flash-Lite), September 24, 2026, one run each.

Where the prompts should actually come from

Don't brainstorm the prompt bank yourself. Build it from what real buyers and reviewers actually say, pulled from five sources:

  1. The challenger's own customer language — G2 reviews, case studies, testimonials, interviews. Extract the jobs people say they're actually doing.
  2. The incumbent's user language — Reddit, Hacker News, X, YouTube comments. Extract why people choose it and where they say it falls short.
  3. The challenger's own positioning — homepage, comparison pages, docs. These are hypotheses to test, not evidence that they're true.
  4. The rest of the competitive set — every other real alternative in the category. This keeps you from falsely treating it as a two-player market when a buyer's actual shortlist has five names on it.
  5. Search and community questions — Reddit threads, autocomplete suggestions, forum questions. These are natural-language buying prompts, not marketer-written ones.

Prompts you write from a whiteboard tend to already contain the answer you expect. Prompts pulled from what buyers actually type do not have that problem.

What to hand back isn't a spreadsheet

The output that's actually useful to a founder or product marketer is not a hundred rows of mentions. It's a short document shaped like this:

A mention count tells you if you're in the room. A crossover map tells you which room you actually belong in.

Where this method flips from research into a problem

Everything above is reading public model outputs, public reviews, and public marketing pages to understand a pattern that's already sitting in plain sight. That's ordinary competitive research, no different from reading a competitor's pricing page or their G2 reviews.

It stops being that the moment the goal shifts from understanding the pattern to attacking it. A few concrete lines:

The test itself is neutral. It answers "where does the recommendation flip, and why." What you do with the answer is the part that has an ethics to it.

The version of this I'd actually run

Six unaided discovery prompts. Ten to twelve compound prompts moving from the plain job to the full buyer-realistic scenario. Six staged-maturity prompts. Six company-size prompts. Six persona prompts. Every important one run with and without the third-party-only restriction, across four models, with the "why" follow-up on every recommendation.

That's a few dozen prompts, not a few hundred, and every one of them earns its place because it moves a specific variable. The output is a recommendation map with real question marks replaced by real answers, a documented crossover point, and a short list of where to stop competing entirely. That last list is usually the part nobody wanted to hear, and it's usually the most valuable thing in the whole exercise.

The best resources I used

If you want help running this for your own category, that's the kind of engagement I take on directly rather than handing you a template. Book a call and bring the two products you're trying to map.

✦
Want your own crossover map?

Stop guessing the boundary.
go measure it.

Book a visibility sprint →