Crawled — not indexed" on author pages is almost never a crawl problem

Quick observation after auditing a handful of sites recently — this came up repeatedly across client work at Megrisoft, where we deal with a lot of technical SEO and GEO audits. Most people see author pages stuck in "crawled — not indexed" and start chasing technical issues — crawl budget, internal links, sitemap gaps. Those things matter, but they're rarely the root cause.

More often, it's one of these:

SEO plugin defaulted to noindex on author archives (Yoast, Rank Math, AIOSEO all do this) The page cleared the noindex check but the content itself is too thin for Google to bother — single paragraph bio, paginated post list, nothing else robots.txt blocking /author/ entirely, which most people set years ago and forgot about

The interesting angle, technically, is that Google is treating "low quality" as an indexing decision rather than just a ranking one. It's not demoting the page — it's just not including it at all. That's a different problem than most crawl audits are set up to catch. ProfilePage schema is underused here too. Implemented correctly, it gives Google a structured signal about who the author is — useful for both traditional search and AI Overview citations. It's something we flag consistently in AEO and GEO audits because answer engines weigh author credibility when deciding what to cite.

Anyone else seeing this pattern? Interested in whether there's a clean way to audit this at scale across large multi-author sites.

Link if helpful: https://www.submitshop.com/author-pages-not-indexed

u/MegrisoftLtd — 1 month ago
▲ 1 r/aeo

Hot take: most "LLM SEO" advice is just regular SEO advice wearing a different hat.

I've been going through some research the team at Megrisoft put together on optimizing for AI-generated answers, and one thing stood out immediately — almost all mainstream advice assumes LLMs work like search engines. That there's a ranking mechanism to reverse-engineer, authority signals being checked, and an algorithm being run.

But that's not how these models work at all.

LLMs don't rank content. They don't store URLs or score your domain during training. What they actually learn is co-occurrence — which brands, concepts, and ideas appeared together, how often, and in what context. That association is what determines whether your name appears in an AI-generated response.

A few things that clicked for me:

  • There's a real difference between training influence (what got baked into the model) and runtime retrieval (when the LLM pulls live search results). Most advice conflates the two completely.
  • High domain authority doesn't directly translate. What matters more is whether your brand is genuinely embedded in conversations about your topic — across publications, forums, and citations — not just your own site.
  • The question isn't "how do I rank in ChatGPT?" It's "when this topic comes up anywhere, is my brand part of that conversation?"

It reframes the whole thing away from algorithm-chasing toward something closer to reputation building at scale.

Curious if others have actually seen results from LLM-specific tactics — or if it's mostly repackaged content advice so far.

Full breakdown here: https://www.megrisoft.com/blog/artificial-intelligence/llm-seo-false-premise-what-to-optimize-for

u/MegrisoftLtd — 4 months ago