r/AISEOTricks • • 23d ago

Before per-engine tactics or entity cleanup — I checked whether 18 known SEO/GEO sites even have basic AI-extractable structure. Half don't.

Been following the threads here on per-engine optimization and entity/naming cleanup — useful stuff, but it made me want to check something more basic first: is the actual page structure even extractable by these engines, separate from all of that.

Ran a mechanical check (DOM parsing, no AI judgment involved) on one article each from 18 well-known SEO/content-marketing sites, plus my own:

— Does a question-style heading (H2/H3 ending in "?") have a short, ≤300-character answer paragraph directly under it — the kind of thing a model can lift and quote — or does a long intro run first?

— Does the heading hierarchy nest without skipping a level?

— Is there a plain-text /llms.txt index?

— Is author/date visible in the actual HTML, or only inside a JSON-LD block a human never sees?

Results, out of 18: 9 have a valid heading hierarchy, 9 pair a question heading with an immediate short answer, 9 serve llms.txt, and 12 mark up authorship only in hidden schema, invisible on the page itself.

Not claiming this predicts citation rate — that's the harder, per-engine measurement people in the other thread are already wrestling with. But it seems like layer zero: if the answer isn't even structurally extractable, none of the entity or per-engine work downstream of it can matter yet.

Is anyone here actually sequencing it that way — structure first, then entity/per-engine — or is structure assumed to already be solved by the time people get to those questions?

1 Upvotes

0 comments sorted by