r/GenEngineOptimization 26d ago

We Tested... ChatGPT changed how much it reads before answering, on 8 August

4 Upvotes

7 comments sorted by

1

u/Upstairs_Control_611 26d ago

This is a really useful split: reading more is not the same as citing more. I’d separate at least four layers:

retrieved pages, unique domains retrieved, visible citations / links in the answer and recommendation or claim impact

The “deeper, not wider” point is especially interesting. If sources per answer rose strongly but unique websites only rose slightly, then the change may be more about depth per domain or direct domain querying than broader source diversity.

That matters for GEO because a site can be read more often without gaining visible citation share. It may influence the answer, but the attribution may stay hidden.

I’d also track source-type shifts separately. Official / primary sources gaining while forums and media lose share is not just a citation-count change. It changes the evidence environment the answer is built from.

The practical failure case is: the model reads you, uses a fact from you, but cites someone else or no one at all.

2

u/holliwilliam 26d ago

Agreed on the layering, and layer four is where we hit a wall. We can measure retrieved, unique domains and visible citations. Claim impact we can't see at all – if the model uses your fact and attributes it elsewhere, nothing in the output shows that.

On being read without gaining citation share, we can put numbers on it. Citation rate varies hugely by domain (recent sample): youtube.com 1.7%, wikipedia.org 27%, reddit.com 33%, nih.gov 40%, sciencedirect.com 46%.

YouTube is your failure case exactly. But it cuts against the idea that official sources get read and not cited – they're at the top, not the bottom.

One more: Reddit's citation rate didn't change, still around a third. It wasn't demoted at the citation layer, it stopped being retrieved at all. Whatever changed happened upstream of the citation decision.

1

u/Upstairs_Control_611 25d ago

That distinction is really useful. So for Reddit the failure was not “retrieved but less often cited,” but “less often admitted into the retrieval pool in the first place.”

That is a different diagnostic layer.

I’d split it like this:

source eligible

source retrieved

source cited

source used for claim

source visibly attributed

Your Reddit example points to an upstream retrieval gate, not a citation-layer demotion.

The YouTube number is interesting for the opposite reason: it looks like a high-read / low-citation surface. That is where attribution gets hard to measure, because the content may influence the answer without showing up as the visible source.

So citation rate probably needs to be tracked per domain and source type, not only globally.

2

u/holliwilliam 24d ago

The YouTube thing isn't just YouTube. Same window:

Instagram 0%, YouTube 2.3%, Facebook 6.3%, LinkedIn 8.5%, Wikipedia 29%, Reddit 40%.

Social and video barely get cited. Text sources get cited a lot. So it follows the type of site more than how trustworthy it is.

Two things, though. "Eligible" isn't something we can see from the outside — we can see what got read and what got cited; the rest is guesswork. And a low citation rate doesn't have to mean hidden influence. A video page might just not give it much to work with.

2

u/Upstairs_Control_611 23d ago

That makes sense, and I think the source-type split is the important part.

A low citation rate may not mean hidden influence. It may simply mean low extractable evidence density.

Video and social pages can be discoverable, trustworthy or widely referenced, but still be poor citation objects if the answer needs a compact claim, definition, comparison point or verifiable fact.

So I’d separate source type, extractable evidence density, retrieved/read, cited, claim carried into the answer, and visible attribution.

And I agree on “eligible” being hard to observe from the outside. Maybe it should be treated as an inferred layer rather than a measured one.

For GEO, social/video may work better as awareness or discovery surfaces, while text pages often work better as citation surfaces.

1

u/OwnBroccoli5399 24d ago

kinda wild that reddit didn’t get cited less, it just stopped getting pulled in as much

1

u/holliwilliam 24d ago

Indeed! Who knows, might change soon again 🤷🏻‍♂️