I’m trying to understand a crawl/indexing issue that has been going on since January.
We run an established site with real organic traffic, around ~3M clicks in the last 90 days, and there are no manual actions in Search Console. The site publishes new landing pages regularly.
The issue is very specific:
Older pages are still crawled and indexed normally. Roughly 99% of our older URLs are fine.
But new landing pages created since January mostly end up in:(Discovered – currently not indexed)And they stay there for months.
For the newer URLs, less than ~20% seem to get crawled/indexed automatically. The rest are discovered by Google, but Google just never comes back to crawl them.The strange part is that if I use URL Inspection and manually request indexing, the same URLs usually get crawled and indexed within minutes. This happens consistently.
So it doesn’t look like Google *can’t* crawl the pages. The pages are accessible, indexable, and apparently acceptable once submitted manually. Google just doesn’t seem to want to crawl them through normal discovery anymore.
We also have another site in the same business/niche, with a broadly similar operating model, page type, and internal-linking strategy. However, the indexing performance is completely different.
One site currently has very poor indexing for new pages, although it used to perform much better in the past. The other site still has a very high indexing rate.
One important difference: one site is structured as a subdirectory, while the other is structured as a subdomain. I’m not sure if that matters here, but I wanted to mention it.
Things we’ve already tried:
- Resubmitted XML sitemaps
- Added stronger internal links to the new pages
- Placed links to new pages in high-visibility areas, including the header, search/discovery areas, and some of the most frequently crawled pages on the site
- Removed/410’d low-value junk URLs to reduce crawl waste
- Cleaned up thin or duplicate pages
- Checked for obvious technical blockers like noindex, robots.txt, canonicals, redirects, etc.
- Compared it against the other similar site we own, where the architecture/internal-linking strategy is broadly similar but the indexing behavior is much better
- Used manual indexing, which works, but the daily limit can’t keep up with our publishing volume
My questions:
- Why would Google crawl and index a new URL almost immediately after manual submission, but not crawl/index it when the same URL is discovered through sitemaps or internal links?
- What signals would make Google “discover” new URLs but decide not to crawl them automatically for months, even on an established site with strong traffic?
- What could cause two similar sites in the same niche, with similar operating methods and page architecture, to have very different indexing rates for new pages, especially when the weaker site used to index much better before?
- Could subdirectory vs subdomain structure affect this kind of crawl/indexing behavior, or is it more likely a site-level quality/trust/crawl-priority issue?
Has anyone seen this pattern and recovered from it without manually submitting every URL?