The situation
Oyster sells into markets where buyers do most of their research before anyone talks to sales. More and more of that research happens inside ChatGPT, Perplexity and AI Overviews, not on a page of blue links. The content library was the front door.
The library itself was not the problem. It was substantial, well written and reviewed by people who knew the subject. The problem was that nobody could draw a line from a published page to an AI citation to a qualified conversation. Without that line, content was a cost center with good intentions attached.
What was in the way
Five things stood between the library and a provable growth channel.
- No baseline
AI visibility lived in anecdotes. Someone would ask a model a question, see the brand or not, and form an impression. No tracked buyer prompts, no citation rate by topic, no way to tell if this quarter beat last quarter.
- Facts in too many places
Product details, legal phrasing and positioning were scattered across docs, old articles and reviewers' heads. When a product was sunset, stale references were even hardcoded into the content generation prompts, quietly reintroducing a deprecated product into new drafts.
- Review was the bottleneck
Legal, brand and editorial all had real stakes in every page, and that should not change. But reviewers were being handed early drafts and asked to do structural work, the slowest and least valuable use of their time.
- Pages aged silently
Published meant finished. Nothing flagged a page as stale, so facts drifted while competitors updated. No trigger ever sent anyone back.
- Technical debt capped the ceiling
Duplicate and missing H1s, thin meta descriptions across whole collections, localized pages that were not actually localized. None of it dramatic alone. Together it limited what any content improvement could return.
The bet: measure before you publish
The instinct in this situation is to publish more. We did the opposite first.
The bet was that a measurement layer would change the roadmap, not just report on it. So the sequence was measurement, then diagnosis, then production. Volume came last, once we knew where volume would pay.
Once you can see citation rate by topic and by page, next quarter's priorities stop being a debate and become a sort order.
That sequencing produced the finding the whole engagement turned on. The fastest gains were not on pages that did not exist yet. They were on pages already published, already carrying authority, and quietly underperforming against buyer questions they should have owned.
How we did it
- Build the tracked prompt set
We started with what buyers actually ask, not what a keyword tool suggests: a commercial-intent prompt template covering the full evaluation journey, run across every priority market. Fragmented tracking setups were consolidated into one coverage map, so we could see which markets and buying stages were monitored and which were blind spots.
- Decompose the prompts
One buyer question does not produce one model query. AI search fans a prompt out into many sub-queries before it assembles an answer. We pulled that fan-out data to see the questions behind the question. That is where most of the citable opportunity lives, and it is invisible if you only track the surface prompt.
- Set the baseline and score by topic
With prompts and sub-queries tracked, citation rate became a real metric, scored per page, per market and per topic. It immediately surfaced a core commercial topic sitting near a 5% citation rate. A topic central to why buyers choose Oyster, and the brand was almost never the answer. That gap became a scoped workstream instead of a hunch.
- Consolidate the source of truth
Before touching a page, we cleaned and consolidated product, brand, legal and positioning facts into one knowledge base, wired into every workflow that touches a fact. The least glamorous step and the highest-leverage one: it is the difference between a positioning change being one operation and forty.
- Design the refresh method
A refresh is not a rewrite. The work was structural: restructure pages so a model can extract a clean answer, front-load the direct answer before the context, map each section to one question, expand FAQ and schema coverage so the page is machine-parseable block by block, and ground every claim in the source of truth. Then push it through the review gates.
- Rebuild production as a pipeline
Every page now moves through the same automated sequence: competitive analysis, outline, research-informed draft, quality audit, tracked-edit rewrite. The source of truth is injected at both outline and draft, so a page reaches a reviewer already fact-checked and on voice. Reviewers get a near-final draft with tracked changes, not a blank page. Their judgment stays where it belongs, and assembly stops waiting on it.
- Measure honestly
Every refreshed page was compared against a symmetric before-and-after window anchored to its own update date, not a shared calendar cutoff. That controls for publish timing and seasonality, so the result survives someone pulling on it. It also meant we saw the losses.
- Close the loop
Every published page enters a ledger with its publish date, performance and a refresh trigger attached. Prompt analytics run continuously. Attribution ties the chain back to qualified pipeline, so the question stops being "does content work?" and becomes "which content do we touch next?"
The results
Across the refreshed cohort, each page measured against its own matched pre-refresh window:
Both bars on one scale. Windows are symmetric and anchored to each page's own update date.
| Metric | Result |
|---|---|
| AI citations | +1,032% (11×) |
| Organic clicks | +110% |
| Refreshed pages that moved up in Google | ~2 in 3 |
| Refreshed pages that gained AI citations | ~78% |
| Per-page citation lift, range | 3× to 50×+ |
The click number matters as much as the citation number. The same structural changes that made pages easier for language models to cite also moved conventional search. AEO and SEO were not two programs competing for one budget. They were the same work.
A small number of pages went flat or slipped. That is the honest shape of any real test, and exactly why the ledger exists. Without a measurement layer those pages would have looked identical to the winners.
Alongside the refresh, the same system handled a full product deprecation cleanup, sitewide phrasing corrections, technical SEO remediation across top library pages, localization fixes, CRO testing on library templates, and a significant expansion of FAQ and schema coverage.
What got installed: the lift is the headline, the system is the asset
Product, brand and legal facts in one place, injected into generation instead of remembered by people. Positioning changes propagate in one operation.
Research to review-ready draft, without re-solving the same problems for every new piece.
Nothing ages silently. The system says what to revisit and when, instead of waiting for someone to notice a decline.
Phrasing, deprecation and compliance changes run as a single pass, not a manual find-and-replace across a large site.
Prompt coverage across priority markets, citation rate by topic, and a traceable path from content to qualified pipeline.
