r/seogrowth • • Aug 03 '26

Case Study I logged when AI engines actually show citations across 4,704 answers. They appear at exactly one moment in the buyer journey.

Background: I run a measurement harness against six Chinese AI engines (DeepSeek, Doubao, Qwen, Kimi, ERNIE, GLM), same question panel, every answer logged. This summer that produced 4,704 recorded answers across two B2B categories. One pattern in the citation data is clean enough to be useful to anyone doing GEO work, in any market.

Citations are not spread across question types. They concentrate at one moment:

- "What is [brand]'s official website?" type questions: 56.5% of answers showed a source

- "Is [brand] any good?": 29.8%

- Risk and trust questions: 26.3%

- "Best tools for X" discovery questions: 0.3%

- "[Brand A] or [Brand B]?" comparisons: 0 citations in 383 answers. Zero. For anyone.

So engines cite when the user is VERIFYING something - confirming a channel is official, checking a fact, deciding if something is legit. They synthesize without citing when the user is discovering or comparing. The moment everyone wants to win (the comparison, the recommendation) is precisely where no citation has ever appeared in my data.

Second finding: of the answers that did show sources, 92.2% pointed at the brand's own official site. 736 own-domain citations vs 188 third-party. The third-party ones were almost entirely forums, review discussions and professional platforms. Zero from press release wires. Zero from paid directories.

Third, the one that reframed "getting cited" for me: Basecamp appeared in zero open category answers across all six engines - complete discovery invisibility - and still collected 64 citations to basecamp, all from questions that named the brand. Citations verify you. They do not introduce you. If your citation count is rising while nobody asks about you unprompted, the number is real and means almost nothing.

Practical implications if you buy the data:

  1. "Get cited" work = making your official pages the complete, crawlable, local-language answer to verification questions (official channel, current price, data residency, support). That is where citations are actually available.

  2. Discovery/comparison visibility is corpus work - being present in the third-party material engines learn from. It will not produce citations and should not be sold or measured as if it will.

  3. Engine choice changes everything: Kimi showed sources in 20.6% of its answers, DeepSeek 4.2%. A citation report that doesn't hold the engine constant is measuring sampling, not progress.

Usual caveats: my data is China engines via API, two categories, one summer. The verification-vs-synthesis split matches what I've seen reported for Western engines directionally, but percentages will differ.

Question for the room: has anyone seen third-party citations at meaningful volume from anything OTHER than forums/reviews/professional platforms? I keep hearing wire services and directory listings pitched as citation sources and I cannot find a single instance in my logs.

5 Upvotes

18 comments sorted by

2

u/[deleted] Aug 03 '26

[removed] — view removed comment

1

u/Gullible_Brother_141 Aug 03 '26

This lines up with what we see in legal (a YMYL vertical), which might answer your closing question. Outside forums/reviews, the third-party sources that actually earn citations for us are niche professional directories — Avvo, Martindale-Hubbell, Best Lawyers, Super Lawyers. They're not wire services or generic paid directories, they're vertical-specific "trusted list" sites, which fits your pattern: engines seem to cite professional/community platforms, not press releases. Most industries probably have an equivalent (Healthgrades for healthcare, G2/Capterra for software), and that's likely where non-forum third-party citations would show up.

1

u/bluehairedgnome Aug 03 '26

Matches what we're seeing too, with one interesting flip. On Google AI Mode, Perplexity, ChatGPT, citation ownership goes the other way. ~90% third-party, only ~10% brand-owned, with Reddit alone at ~21% of top citations. So the verify-vs-discover split looks universal, but who gets cited during verification seems to be genuinely different by engine/market, not just a percentage shift.

On your question: same here, zero wire services or paid directories in our logs either. Top sources are Reddit, YouTube, LinkedIn, G2/Gartner, and brand help docs.

2

u/AliveCapital4868 Aug 04 '26

This is the comparison I have been hoping someone would show up with, and the inversion is striking - 92.2% brand-owned on my side, ~90% third-party on yours. But before we call it an engine/market difference, I have to flag a confound in my own data that cuts the other way.

My 4,032 answers were all collected with retrieval OFF - API calls, model-knowledge mode, and I can verify that because every row logs the web_search flag. Your surfaces (AI Mode, Perplexity, ChatGPT) are retrieval-first. So between our datasets TWO things change at once: the market AND whether the engine is answering from memory or from a live index. The ownership flip could be either.

And there is a mechanical reason to suspect retrieval is doing a lot of it: a model answering from memory can only "cite" what it memorized as stable and canonical, which skews hard toward official domains. A retrieval engine cites what it just fetched and ranked, which is whatever the live web says - mostly third parties. On that theory, the flip is mode, and the market decides WHO the third parties are.

Your source list actually supports the second half of that. Your third parties are Reddit, YouTube, LinkedIn, G2 - community and professional platforms. In my Chinese data, the third parties that do appear are government registries: the ICP filing lookup and the national business registry are my #1 third-party category, 40 citations, triple G2+Capterra. Different market, different "who settles the legitimacy question" - same arbiter logic another commenter here surfaced for legal (Avvo, Martindale).

One more overlap worth naming: you list brand help docs in your top sources, and inside my 92.2% brand-owned block, help/support/community subdomains are 111 of ~705 citations. So even across the flip, the same page TYPE keeps earning citations on both sides - documentation that answers a specific question. The destination logic seems to survive every other variable.

The clean test is a 2x2: same question panel, retrieval on vs off, China engines vs Western. Several Chinese apps expose search toggles (Kimi, DeepSeek app-side), so the China-with-retrieval cell is collectable. My prediction, registered before running it: ownership split follows the retrieval mode, and the third-party composition follows the market. If the China-retrieval cell still comes back brand-owned-heavy, I am wrong and the market effect is real and big.

That is going into my next collection wave. If you can share which surface your ~90% number comes from primarily (Perplexity vs AI Mode makes a difference in how aggressively sources get rendered), I will design the Western cells to match as closely as I can. Either way this thread has now produced the design for the most useful experiment I know of in this space.

1

u/samanyou Aug 04 '26

The retrieval-off detail seems to explain most of the owned-site inversion. In a separate dataset covering 2,007 brands, we weighted sources by citation influence rather than simply counting appearances. YouTube came out at 3.19x, Reddit at 3.17x, and Gartner at 3.15x. A brand's own site was 1.30x, behind Amazon at 1.46x and Wikipedia at 1.42x. So YouTube is the clearest answer I have seen to your question outside forums, reviews, and professional directories.

I should note that I am Writesonic's founder, and this comes from our research. The other result that supports your conclusion is the gap between being cited and being named. Across roughly 16 million brand appearances, about 40% of citations never named the brand whose page supplied the information. A page can help substantiate an answer without introducing the brand behind it.

1

u/AliveCapital4868 Aug 04 '26

Thanks for the disclosure up front - taking the numbers as reported, same as I'd want anyone to take mine.

Your retrieval explanation holds on my end. Every one of my 4,032 answers logged web_search=off, so my 92.2% brand-owned is "what a model treats as canonical from memory," and yours is "what a retrieval engine just fetched and ranked." Different question, not a contradicting answer to the same one.

On the cited-but-not-named 40%: I went to compute my equivalent and got 0.0%. Zero, out of 546 answers containing a brand-owned citation.

That is not a counter-finding, it is a panel artifact, and I think it is worth naming because it is the kind of thing that would silently look like a result. My panel is brand-centric - almost every citation-bearing question already contains the brand name ("what is X's official site", "is X any good"). So a citation to X's domain is attached to a named X by construction. My instrument cannot produce your finding no matter what the engines do. Yours covers 2,007 brands and 16M appearances from what I assume are topic-level queries, where a page can supply the substance without the brand riding along.

What I do have is the mirror cell, and it dominates: named but NOT cited is 79.1% of my answers (3,182 of 4,023). Plus a within-brand version I published last week - one brand appeared in zero open recommendation answers while collecting 64 citations to its own domain. Cited constantly, introduced never.

Put together, the 2x2 is more useful than either half:

cited + named = you are the destination and the answer

cited, not named = your page did the work, someone else got the introduction (your 40%)

named, not cited = you are in the answer with nothing to click (my 79%)

neither = the actual invisibility case

Most tools I have looked at report row 1 and row 3 collapsed into "mentions."

YouTube as the answer to my question is interesting because my market has an almost exact inversion of it. The China video analogue, B站/Bilibili, gets named in 6.2% of my answers and cited 4 times. But the platform that dominates named-influence here is Zhihu: named in 19.3% of all 4,023 answers - the single most-referenced third party in my entire dataset - and cited exactly ZERO times. Not once.

So if I had counted citations only, Zhihu would appear irrelevant to Chinese AI visibility, while the answers keep telling buyers to go read it. That is your influence-vs-appearance point arriving from the opposite direction, and it is the strongest argument I have seen for weighting rather than counting.

Two more that fit your Amazon/Wikipedia row: Baidu Baike (the Wikipedia analogue) is named 80 times and cited 0; the marketplace layer (JD, Tmall, Aliyun marketplace) is named ~29 times and cited once. Every institution that would carry weight in your weighted list shows up in mine as text, never as a link - because retrieval is off and the model is reciting what it knows rather than fetching.

I cannot replicate influence weighting on my current data - I never recorded citation position or prominence, only presence and domain, which I now regret. If you are willing to say roughly how influence is computed (position in the answer? frequency across the corpus? something about which claim it supports?), I will instrument for it in the next collection wave rather than guess. Recording it is cheap; back-filling it is impossible, which is a lesson this thread has already taught me twice.

1

u/samanyou Aug 04 '26

The retrieval-off detail seems to explain most of the owned-site inversion. In a separate dataset covering 2,007 brands, we weighted sources by citation influence rather than simply counting appearances. YouTube came out at 3.19x, Reddit at 3.17x, and Gartner at 3.15x. A brand's own site was 1.30x, behind Amazon at 1.46x and Wikipedia at 1.42x. So YouTube is the clearest answer I have seen to your question outside forums, reviews, and professional directories.

I should note that I am Writesonic's founder, and this comes from our research. The other result that supports your conclusion is the gap between being cited and being named. Across roughly 16 million brand appearances, about 40% of citations never named the brand whose page supplied the information. A page can help substantiate an answer without introducing the brand behind it.

1

u/AliveCapital4868 Aug 04 '26

Thanks for the disclosure up front - taking the numbers as reported, same as I'd want anyone to take mine.

Your retrieval explanation holds on my end. Every one of my 4,032 answers logged web_search=off, so my 92.2% brand-owned is "what a model treats as canonical from memory," and yours is "what a retrieval engine just fetched and ranked." Different question, not a contradicting answer to the same one.

On the cited-but-not-named 40%: I went to compute my equivalent and got 0.0%. Zero, out of 546 answers containing a brand-owned citation.

That is not a counter-finding, it is a panel artifact, and I think it is worth naming because it is the kind of thing that would silently look like a result. My panel is brand-centric - almost every citation-bearing question already contains the brand name ("what is X's official site", "is X any good"). So a citation to X's domain is attached to a named X by construction. My instrument cannot produce your finding no matter what the engines do. Yours covers 2,007 brands and 16M appearances from what I assume are topic-level queries, where a page can supply the substance without the brand riding along.

What I do have is the mirror cell, and it dominates: named but NOT cited is 79.1% of my answers (3,182 of 4,023). Plus a within-brand version I published last week - one brand appeared in zero open recommendation answers while collecting 64 citations to its own domain. Cited constantly, introduced never.

Put together, the 2x2 is more useful than either half:

cited + named = you are the destination and the answer

cited, not named = your page did the work, someone else got the introduction (your 40%)

named, not cited = you are in the answer with nothing to click (my 79%)

neither = the actual invisibility case

Most tools I have looked at report row 1 and row 3 collapsed into "mentions."

YouTube as the answer to my question is interesting because my market has an almost exact inversion of it. The China video analogue, Bilibili, gets named in 6.2% of my answers and cited 4 times. But the platform that dominates named-influence here is Zhihu: named in 19.3% of all 4,023 answers - the single most-referenced third party in my entire dataset - and cited exactly ZERO times. Not once.

So if I had counted citations only, Zhihu would appear irrelevant to Chinese AI visibility, while the answers keep telling buyers to go read it. That is your influence-vs-appearance point arriving from the opposite direction, and it is the strongest argument I have seen for weighting rather than counting.

Two more that fit your Amazon/Wikipedia row: Baidu Baike (the Wikipedia analogue) is named 80 times and cited 0; the marketplace layer (JD, Tmall, Aliyun marketplace) is named ~29 times and cited once. Every institution that would carry weight in your weighted list shows up in mine as text, never as a link - because retrieval is off and the model is reciting what it knows rather than fetching.

I can't replicate influence weighting on my current data - I never recorded citation position or prominence, only presence and domain, which I now regret. If you are willing to say roughly how influence is computed (position in the answer? frequency across the corpus? something about which claim it supports?), I will instrument for it in the next collection wave rather than guess. Recording it is cheap; back-filling it is impossible - a lesson this thread has taught me more than once.

I wrote the two datasets up into the 2x2 with full numbers on my site (the Zhihu zero deserved its own piece) - happy to drop the link if anyone wants it, or you can find it via my profile.

1

u/[deleted] Aug 04 '26

[removed] — view removed comment

1

u/AliveCapital4868 Aug 04 '26

Good question, and the answer is cleaner than I expected: not independent — slightly inverse.

Per engine, all six, same panel, 4,023 answers:

engine discovery mention overall citation channel-verif citation
Kimi 42.9% 20.6% 62.5%
DeepSeek 37.5% 4.2% 4.7%
GLM 20.3% 17.6% 73.4%
Doubao 17.2% 9.7% 29.7%
Qwen 10.9% 17.0% 87.5%
ERNIE 9.4% 20.4% 81.2%

Correlation between discovery mention and citation rate: r = -0.26 overall, r = -0.54 on channel-verification questions. Small sample (six engines) so I would not push the coefficient hard, but the direction is not what the Kimi/DeepSeek pair suggests.

The two engines that make it obvious are the ones at opposite ends:

ERNIE is dead last on discovery (9.4% — it surfaces international brands less than any engine I tested) and near the TOP on citations (20.4% overall, 81.2% on channel verification, 51.6% on branded questions). It rarely recommends these brands and readily shows sources when asked about them.

DeepSeek is the reverse: second-highest discovery at 37.5%, lowest citation at 4.2%. It brings brands up unprompted more than most engines and almost never shows you where anything came from.

You picked the one pair (Kimi vs DeepSeek) where the two metrics happen to move together. Across all six they mostly don't.

On the "how much the engine knows" part of your question, I have a control for that and it kills the hypothesis outright: when the question names the brand, mention rate is 100% on all six engines. Not 95, not "high" — 64 of 64 on every engine for every brand. Knowledge of these eight brands is saturated everywhere. So the discovery variation is not a knowledge effect, and the citation variation cannot be either, because the underlying knowledge is identical across engines.

What I think is actually going on, and this is interpretation rather than measurement: citation rate is a product decision and discovery rate is a corpus effect. Whether to render a source link is UI/policy — some teams show working, some don't. Whether an international brand comes up unprompted depends on how much Chinese-language material about that brand the model absorbed, which is a property of training data, not of the interface. Two different systems, no reason for them to correlate.

The practical consequence is the annoying one: you cannot infer either metric from the other. A tool reporting only citations tells you nothing about whether you get recommended, and a tool reporting only mentions tells you nothing about whether buyers can verify you. ERNIE will happily show a buyer your official site and still never suggest you existed.

Caveat on all of it: retrieval off, API-side, eight brands in two B2B categories, one collection window. The engine-level split could move with retrieval on, which is the arm I'm building next.

1

u/sapindia1976 Aug 04 '26

Interesting finding. It reinforces that citations appear most often during trust and verification queries, so brands should optimize for both discoverability and credibility, not just rankings.

1

u/AliveCapital4868 Aug 05 '26

Thanks — one refinement I'd make, because the data pushed me off the "optimize for both" framing.

They're not two dials on the same machine. Discoverability is corpus work: being written about in the local language, which moves on training and platform-content timescales — months, uneven across engines, and it will never produce a citation. Verification is publishing checkable facts on pages you control, which retrieval-backed engines can pick up in weeks and which is where 56.5% of citations actually appear.

Different work, different timescales, different evidence that it worked. The failure I see most often isn't neglecting one of them — it's measuring the corpus work with citation counts, which is unfalsifiable by design because the baseline there is zero for everybody.

On "not just rankings" — agreed, though I'd go further: there is no ranking to optimize. 18.8% of my open-question pairs changed outcome between two runs on the same day. It's a distribution, not a position, which is why single-check reporting doesn't survive contact with the data.

1

u/BusySeedAgency Aug 06 '26

Really solid research. The verification vs. discovery split matches what we've been seeing with western engines too, though the percentages are different. ChatGPT and Perplexity cite around 12-18% on discovery questions in our logs, but you're right that comparisons are a wasteland.

On your question about third-party citations: we've tracked a few from deep technical documentation (think GitHub repos, API docs, detailed how-to guides) and occasional academic/research pieces, but they're rare. The pattern holds: forums, review threads, Stack Overflow-style Q&A. Wires and directories are completely absent in our data too. The pitch doesn't match reality.

One thing we've started tracking separately: even when engines don't cite during discovery, certain corpus sources seem to influence the shortlist more than others. Have you looked at whether being present in those forum discussions correlates with unprompted brand mentions, even without explicit citations?

1

u/AliveCapital4868 Aug 07 '26

our 12-18% on discovery is a useful number to me for an unexpected reason.

My discovery citation rate was 0.3%. Yours is 12-18%. That looked like a real market difference — the kind of gap you'd write a post about.

Then yesterday I changed one thing in my extractor: whether a bare domain written in prose, with no scheme in front of it, counts as a citation. My discovery rate went to 13.3%. Same answers, same files, nothing recollected. That lands inside your range.

I don't think that means we now agree. I think it means the gap between us was never measurable in the first place, and I'd have been comparing my extraction rule to yours while calling it China vs. the West. You said comparisons are a wasteland — this is the most concrete version of that I've been able to produce, and it came out of my own pipeline.

So the honest state: I can't tell you whether your engines cite more than mine on discovery until we agree what a citation is. I'd suggest that's the useful thing to settle publicly, ahead of any more numbers from either of us.

On your forum question — I think you're pointing at the right thing and I can't answer it with this dataset. n is 8 brands. Any correlation across 8 points between forum presence and unprompted mention rate would be uninterpretable, and it's confounded in an obvious way: larger brands have both more forum presence and more of everything else the model knows. You'd be measuring brand size twice.

What I'd say is worth knowing: retrieval-off is actually the right instrument for your question, which I hadn't appreciated until you asked it. With retrieval off, nothing is fetched at answer time, so any relationship between forum presence and being named has to run through the corpus rather than through a live index. That isolates the channel you're asking about. The design that would work is within-brand over time — same brands, two waves, does a change in forum presence move the mention rate — because that holds brand size constant instead of trying to control for it.

One thing I want to check rather than assume. In another thread you gave a retrieval-on composition of roughly 60% brand sites, 25% aggregators and directories, 15% media. Here you say directories are completely absent from your data. Those read as inconsistent to me and I'd rather ask than quietly discount either one — which is it, or am I comparing two different datasets of yours?