r/TechSEO

▲ 6 r/TechSEO+1 crossposts

What if “SEO-friendly” URLs are actually a disadvantage for AI crawler discovery?

SEOs have been taught for decades that URLs should be descriptive. A URL such as /wordpress-performance-optimization/ is considered better than something meaningless like /x7ab31/.

But that assumes the crawler actually needs to fetch the page before deciding what it is about.

I've been watching AI crawler behavior more closely, particularly GPTBot and ClaudeBot, and something made me question that assumption.

If a crawler can infer enough about the likely content from the URL, link context and surrounding semantics, it can also decide that the page isn't worth fetching - without ever seeing the actual content.

So I tested the opposite approach.

I exposed alternative Markdown representations of existing content through completely opaque URLs. The URLs contained no topic, keyword or other clue about what was behind them.

Both GPTBot and ClaudeBot discovered and fetched all of them.

This obviously doesn't prove that the content is used for training, citations or AI answers. It only proves something much earlier in the chain: the content was actually retrieved.

And that made me wonder whether GEO is currently starting too late.

Most GEO advice focuses on optimizing content for AI systems after discovery: structure, entities, concise answers, citations, schema, etc.

But what if the first optimization layer should be:

Make sure the AI crawler actually sees the content in the first place.

In that context, a descriptive URL may not always be an advantage. It may also give a crawler enough information to reject the content before fetching it.

Maybe AI discovery needs to be treated as its own optimization layer, separate from traditional SEO.

Has anyone else observed crawler selection happening before the actual page fetch?

reddit.com
u/Good_Flight6250 — 1 day ago

Title: Would you use a temporary sitemap to speed up the deindexing of 300 URLs?

I have a relatively small website with around 600 indexed pages, and I want to remove roughly 300 of them from Google.

I've already applied noindex/removed the relevant pages, but the process has been pretty slow. After about a month, only around 50 URLs have disappeared.

I'm looking at different ways to speed things up.

One idea I had was to temporarily create a separate sitemap containing the 300 URLs I want deindexed, hoping that Googlebot would recrawl them faster and detect the noindex or 404/410 response.

Once they were deindexed, I would remove that sitemap.

Has anyone tried something similar?

Do you think a temporary sitemap containing URLs you want removed could help speed up recrawling, or would you simply remove them from the regular sitemap and wait?

I'm also considering using GSC Removals to speed up their disappearance from the SERPs, although I understand that's only temporary and I'd still need to keep the noindex/404/410 in place.

I'm especially interested in real-world experiences with hundreds of URLs rather than millions. What strategy worked best for you?

reddit.com
u/TheChiuaua — 22 hours ago

Datacenter proxies for ecommerce price monitoring, sessions keep dying mid scrape...

We’re running a small ecommerce monitoring tool. The goal is to check competitor prices/stock every hour across like 40 product pages. We've tried datacenter proxies from some random cheap host and sessions just drop halfway through. It sucks cause then I gotta restart the whole scraper. We need something where the session stays alive long enough to finish a batch for sure. ISP proxies keep coming up as an option too but idk if theyre actually more stable or if it's just fancy marketing. Curious to hear whats actually holding up for you guys for this kind of monitoring. What do you guys use or how do you even scrape?

reddit.com
u/bg81011 — 1 day ago
▲ 52 r/TechSEO

Sudden Google Search Console impressions drop

Hi everyone,

I noticed a very sudden drop in my Google Search Console performance.

My site was getting around 12,000 impressions per day, but suddenly dropped to almost 0 impressions, with only a very small number of clicks.

What’s confusing is that all of my pages are still indexed in Google Search Console, and I don't see any obvious indexing issues.

I haven't intentionally made any major SEO changes before the drop.

Has anyone experienced something similar? Is this something that can happen temporarily with Google, or should I be looking for a specific issue such as a manual action, algorithm update, indexing problem, technical SEO issue, etc.?

What would you check first in this situation?

Thanks!

u/deniscas09 — 2 days ago

Help: Showing up in Bing ≠ Being indexed in Bing

I noticed across several of my websites that 90% of the pages are indexed in Bing (submitted in Bing Webmaster Tool and reported as such in Bing Site Explorer) but when I actually search for a lot of them in Bing, they don't show.

I heard Bing has its own filter and may or may not show an indexed page in the SERP.

Does anyone have tips (e.g. technical SEO or else) to maximise the chances indexed pages do show in the SERP?

reddit.com
u/phb71 — 1 day ago

Is google SandBox real?

I am not selling anything, actually I am pretty dumb when it comes to SEO.

I launched my site somewhere in match, and it was climbing steady, no big deal (20 - 50 clicks a day and ~4-6k impressions). But, starting July 8, I started to get pretty much DOUBLE (100 clicks and around 11k impressions).

Bu that's not all, I was looking at my phone in GA4 analytics and from August 1, I was having "Realtime users" like 300 persons in the last 30 minutes.

Now, I am getting stable 80k impressions and around 1.5k clicks from google daily.

NOTE:

  1. I didn't change A THING on the site.

  2. I am publishing the same 1-2 articles a week.

  3. I have 2 backlinks (someone random linked to me).

Now to the question: I was reading in subreddits and I found this google sandbox (like you wait 6 months and then you see real traction, or you flop). Is this even real? Maybe there is some technical stuff that I am not aware? If not, how comes I get PER DAY more clicks and impressions than I was getting per month literally 30 days ago? OR maybe I am just lucky?

u/Dota2ProTips — 2 days ago
▲ 31 r/TechSEO+3 crossposts

I accidentally built a very cool automatic SEO tool... so take it and help me improve it.

I figured I would post this here, and some of you would get a kick out of it. I have been working on my personal site, and it has grown considerably. I notoriously suck at SEO, so I built a tool that can help me with it. It is still early, but proven in production on my site (and now several others). The next version will include refinements, expanded MCP access, and AEO/GEO features.

If anyone wants to kick the tires and give me feedback to improve it, I would appreciate it. It is free, open-source, MIT-licensed, and I would like to make it better for myself and others.

https://github.com/awizemann/seo-agent

u/awizemann — 2 days ago

Migration aftermath: 90% of 2 domains out of 22 "Crawled - currently not indexed" for 10+ months

Hi, I already went through all the "Crawled - currently not indexed" posts here, but in my mind this situation is a bit more specific, so I would really appreciate your input.

  • September '25: we migrated our 22 domains (ecommerce).
  • November '25: 20 domains migrated without issues, while 2 of them got only 2-3 pages indexed, the remaining got stuck in "Crawled - currently not indexed".
  • Present days: those 2 domains got 5-10 pages indexed. As a result, I'm losing my mind 🥲

Context: the site content and architecture is not great across all 22 domains. I took over SEO only recently, and I've been comparing all the comparable (authority, link profile, code, content, rebranding campaigns) - nothing sets apart these 2 faulty domains from the other 22. Furthermore, one of these 2 domains is a big market/high traffic, and the other is a small market/small traffic, so I can't even go "Well, they never got traffic even before the migration" (additional mind lost).

OF COURSE, old domains for these 2 are still greatly indexed 🚑 it really looks like G is ignoring the 301s (correctly set 301s, set the same way across the other domains). DuckDuckGo/Bing index these 2 domains without issues.

Have you ever seen anything like this?

Has anything fixed it, or do you have any tips?

Could it be G is more annoying in some countries than others?

I've managed other migrations previously in my career, and never ever happened. For completion of info:

  • change of address in GSC was done properly;
  • none of the 2 domains had traffic drop before migration (while others that had traffic drop, indexed without issues);
  • I'm regularly pushing indexing for high potential pages, but nada. I'm fixing content, internal linking and site structure, but apart from those additional 4-8 pages indexed, nada;
  • Products are visible in G freelisting (only source of traffic for these 2 domains 🥲 ), so it knows people want them and click them, still actively decide to not index them.

Thanks in advance!

(uplifting memes also accepted)

reddit.com
u/Any-Leading9879 — 2 days ago

Merging site B into A, then renaming A to B's domain. Stage them or do both at once?

Setup: Two sites in the same category, overlapping product lines but not identical. A is larger; B is smaller but owns the name you'd actually want. Merging B into A, then renaming A to B's domain.

Plan: Build a complete URL list. Use rank data to pick winners where A and B overlap. Merge the content and rename the domain in a single cutover, every old URL single-hopping to its final home.

Question: would you stage these, or do them together? Staging means B's URLs move twice. Doing them together means two changes on the same day. Wanna talk me out of the latter?

reddit.com
u/uncoolcentral — 1 day ago

Share what you're working on (including what you're building)

We want to support creators, but we had to enforce the no shilling rule because it was getting out of hand. You now have a weekly thread.

This is the one place you can shill for your products, ask for feedback, etc. Keep it here or you risk being banned. And keep it related to technical SEO.

reddit.com
u/AutoModerator — 3 days ago

Old /en/ URLs still have 88% of impressions three months after the migration

I moved a small site from /en/* URLs to root-level URLs about three months ago. I thought the migration was mostly settled until I exported the page report from GSC and actually added up the old URLs.
In the May 10–August 9 window, 97 retired /en/* URLs had 8,821 impressions and 25 clicks. The whole site had 9,986 impressions and 39 clicks. So the old URLs were still getting 88.3% of impressions and 64.1% of clicks.
One old URL by itself had 908 impressions and 20 clicks. That URL had not served content for months, but it accounted for just over half of all clicks on the site during the period.
I went back through the migration because my first assumption was that I had missed something obvious. The old URLs return direct 308s, with no chains. They are not blocked by robots.txt. The sitemap and internal links only contain the new URLs, the new pages are self-canonical, and there is no old hreflang setup left.
It was not a completely clean sweep. I found seven old URLs redirecting to destinations that now returned 404. That was my mistake, and I fixed those separately. They had about 58 impressions between them, so they do not explain what is happening with the other 90 URLs.
What confused me most was this:
• August 3 export: 7,279 impressions on old URLs (94.7%)
• August 11 export: 8,821 (88.3%)
• August 12 export: 9,028 (86.1%)
These are rolling three-month exports, not identical date ranges, so I know this is not a clean comparison. But I had already looked at the falling percentage and told myself consolidation was improving. Then I noticed the absolute number was going up. Most of the percentage change came from new content increasing the site total.
The latest seven-day window I have, August 4–10, still shows 45 old URLs receiving 1,326 of 2,183 impressions.
My current suspicion is simply that many of the old URLs have not been recrawled often enough. I removed them from the sitemap and replaced every internal link during cleanup, which also means there are fewer current signals leading Googlebot back to them. That is only a theory, though. I cannot prove it from the performance report.
I also compared the eight old URLs with the most impressions with their new equivalents. Three new URLs have their own rows in GSC. Five are completely absent from the page report.
I realize that an absent GSC row does not mean zero impressions because of reporting thresholds. Still, I cannot find anything that separates the three that appear from the five that do not. It is not publish date, position, or whether the slug changed.
If you have handled a similar path migration on a small site, did you see both versions reporting for a while? Or did the new URL only begin appearing after the old one dropped out? I am trying to work out whether those five missing new URLs are a normal in-between state or whether I have overlooked another signal.

reddit.com
u/Ivyyy1015 — 3 days ago
▲ 28 r/TechSEO+27 crossposts

Things have been getting crazy for my SaaS recently 🔥

Around 2-3 weeks ago I launched my SaaS, SeoLoupe.

It is a lightweight SEO tool that helps you find and fix SEO issues holding your website back, currently I am at 432 users and 5 paying users.

Essentianly the main purpose is to help your website rank higher on Google search and LLMs.

Since I am always trying my best to improve the product, I am happy to answer any questions or any feedback in the comments.

(here is the product if you want to check it out)

u/megatech_official — 5 days ago

I mechanically checked 18 well-known SEO/content sites for "AI-answer readiness" — including my own. Half fail a basic heading-hierarchy check.

Wrote a small script to run a mechanical check (DOM parsing, no AI judgment involved) against one article each from 18 well-known SEO/content marketing sites, plus my own. Wanted to share the pattern, not pitch anything — no tool name, no link, just the method and the numbers. Happy to describe the script in comments if anyone's curious.

What I checked, one article per domain, one snapshot in time:

  • Does the site serve /llms.txt?
  • Is there a question-style H2/H3 (ends in "?")?
  • Is there a short (≤300 char) answer paragraph immediately under that heading — the kind of thing an AI engine could lift and quote directly, as opposed to a long intro before the real answer?
  • Does the heading hierarchy nest without skipping a level (H2 → H3 → H4, not H2 → H4)?
  • Is author/publish-date visible in the actual HTML, or only inside a JSON-LD block a human would never see?

Targets: Ahrefs, Semrush, Moz, Backlinko, HubSpot, Search Engine Journal, Search Engine Land, Yoast, Kinsta, WP Engine, WPBeginner, SurferSEO, Content Marketing Institute, Animalz, Siege Media, Foundation Inc, Omniscient Digital, and my own site. 18 of 20 targets resolved; 2 didn't (environment issue on my end, not excluded for any other reason).

Results, out of 18:

  • 9/18 serve llms.txt
  • 10/18 have a question heading somewhere
  • 8/18 pair that heading with an immediate short answer
  • 9/18 have a valid heading hierarchy (the other 9 skip a level or start below H2)
  • 12/18 mark up author/date only in JSON-LD, not visibly in the page

My own site came out with an invalid heading hierarchy too — the FAQ is correctly marked up in schema but isn't built from real H2/H3s in the HTML, so it fails the exact same check. Fixing that next.

Not claiming this predicts who gets cited more in ChatGPT/Perplexity — that's a different, harder measurement over time. This is just: even among sites whose whole business is search/content strategy, basic machine-extractable structure is inconsistent, and a lot of the "structured data" people cite as a GEO win is invisible to anything except a schema parser.

Happy to share the raw JSON if anyone wants to poke at the methodology or thinks a check is measuring the wrong thing.

reddit.com
u/Gullible_Brother_141 — 4 days ago
▲ 9 r/TechSEO+2 crossposts

Anyone got a good Data Studio dashboard for tracking Search Console and/or Analytics?

I have been experimenting with creating a Data Studio dashboard that pulls Search Console metrics. There a several free templates online, but they are quite basic.

I am trying to figure out which key metrics to include to monitor client sites quickly and easily. I am moving into an SEO role at an agency, and previously I just monitored my own site through Search Console. That is too time consuming yo poke around at scale with multiple client sites.

Some ideas I have are tricky to implement as a beginner, so it’s trial and error with ChatGPt help.

Does anyone have a good dashboard they have created with Search Console and/or Analytics worth sharing?

reddit.com
u/thinkit_doit — 5 days ago
▲ 45 r/TechSEO+1 crossposts

Lets Debate how Googlebot (aka Spiders) work and Don't work; a google Doc problem

I think Google has done a bad job of breaking down how spiders work and the predominantly shared understanding is that Spiders routinely visit your site/page, are constantly on the look for updates and that Google crawls your site A to Z, Home page - tier by tier and tries to assess and "understand it"

I think its important for people at all levels of SEO to debate and talk about how Googlebot actually works.

Googlebot reality vs the Spider Fairy tale?

Is this a fair description? Feedback below please.

In the world described by Google's "Spider" story,, we see narrative like

  • Bots/Spiders wont index content that is thin, duplicate, doesnt contain new content
    • This is nowhere to be seen in Googles documentation
    • Googlebots dont grade content
  • XML Sitemaps force Google to index pages
    • XML sitemaps aren't read just because they get update
    • Again BIG traffic sites v Little traffic sites
  • Googlebots crawl your whole site and assess it
    • Actually - pages are crawled by "pools"

Claude can piece together the story quite well

  • Googlebot fetches pages for the Indexing service
  • It doesnt process content
  • It cannot grade content
  • All Googlebots are fully chromium
    • Still - we see problems with SSR v CSR

The Problem With Google Docs

What I'm seeing is that the "story" about crawlers - like on Google's "non-technical" how it works pages and videos about bots as spiders are just wrong and misleading.

Whereas the narrative told at Google Search events doesn't align with the Google Dev Guide

Google describes crawling as starting with "URL discovery" — there's no central registry of all web pages, so Google constantly looks for new and updated pages and adds them to its list of known pages; pages are discovered when Google extracts a link from a known page to a new page.

Critically, the docs also state that

>"Google doesn't guarantee that it will crawl, index, or serve your page, even if your page follows the Google Search Essentials."

Telling the Real Story

These snippets of information are all over the web, not in Google's repository

Gary Illyes — Stone Temple Consulting

Q&A (2016): Illyes explained Google's internal "host load" concept, which sets a bucket of URLs in importance order that Googlebot crawls in that order — "that is driven mainly by the importance of the pages on a site but not by the number of URLs." He gave examples: a URL in a sitemap is deemed more important, high PageRank URLs "probably should be crawled more often," and "basically the more important the URL is" the sooner and more often it gets crawled.
Link: https://www.seroundtable.com/googles-gary-illyes-crawl-budget-scheduling-host-load-22097.html

Search Off the Record — crawl scheduling episodes

(Splitt/Illyes): Google's "crawl scheduler" estimates which pages are due for a recrawl and runs "discovery crawls" for sections likely to contain new URLs; the system organizes URLs by assigning a priority influenced by content quality and update frequency.
Link (recap of multiple episodes): https://aminforoutan.com/blog/google-crawlers-insights-search-off-the-record-podcast/
Podcast home: https://search-off-the-record.libsyn.com

u/WebLinkr — 7 days ago

How do you write SEO articles with Claude that actually rank?

What’s your proven workflow for creating SEO content with Claude or other AI tools?

  • What process do you follow from keyword research to the final article?
  • Are your AI-assisted articles ranking consistently in Google?
  • What have you found works best, and what should be avoided?

Would love to hear from people who have tested this at scale.

reddit.com
u/Expensive_Spare821 — 6 days ago
▲ 10 r/TechSEO

Help getting e-commerce indexed

I have an e-commerce site with over 5 million products that only got 17k pages indexed out of 1.9M crawled. Any recommendations on speeding up the index rate?

reddit.com
u/ded-ebar — 7 days ago
▲ 1 r/TechSEO+1 crossposts

How should we actually audit llms.txt? I built a Chrome extension to find out

I’ve been playing with llms.txt audits for a while and ended up building a Chrome extension around it: LLMs.txt Compliance Inspector v1.2.

Important disclaimer before anyone reaches for the pitchforks: I’m not claiming llms.txt improves rankings or AI citations. The evidence for that is weak at best right now.

What interested me was a different problem: if people are going to create these files anyway, how do we evaluate whether they’re actually well structured and useful?

Most checkers I found basically answer:

“Does /llms.txt exist? ✅”

I wanted something closer to a technical audit.

It checks llms.txt and llms-full.txt, structure, links, missing sections, technical issues and recommendations.

I also deliberately treat different sites differently. A Shopify store, WooCommerce store, blog, SaaS/API documentation site and corporate website shouldn’t necessarily get the same checklist.

This is still very much my interpretation of an emerging convention, which is exactly why I’d like TechSEO people to break it.

Chrome extension:
https://chromewebstore.google.com/detail/llmstxt-compliance-inspec/kobabmaojpooanphhogepogipcnajcao

I’m particularly interested in:

• checks you think are bullshit
• checks I’m missing
• things that should be warnings rather than errors
• whether site-type-specific auditing makes sense at all

Roast away 😄

u/uiarago — 6 days ago
▲ 3 r/TechSEO+1 crossposts

Bridge CDN: pre-launch phase has began

So I'm working on something exciting and aiming really high — compete with CloudFlare in CDN, DNS, and edge computing niche, crazy right?

I've built my first CDN years ago around 2012, that time "Content Delivery Network" wasn't even a broadly known thing, niche had two large players -- Brightcove and Akamai, their entry price was 25K+ and was beyond what we needed at that time. Tech news outlets were full of DDoS attack articles (it was cheap and allowed putting your competitor's website down with ease), and Chrome browsers were about to block all websites without SSL/HTTPS with scary warning of "not encrypted connection". That were times when Cloudflare got its spin with free for all SSL-certs (thanks to Let's encrypt foundation).

Nowadays we are facing another major shift in who's the primary user of the web, do you still think it's humans? Wrong, AI-agents, crawlers, and bots are dominating in the web since 2025, -- automated traffic you new primary content consumer.

The main Bridge CDN goal is to bridge the gap between web that was build for humans and its new main consumers. With built-in SEO/AEO/GEO tools and services bots will be able to access your website and its pages within milliseconds and understand its context with ease. All new and changed pages will get sent to search engines via IndexNow protocol for the fastest discovery, and your devs won't need to worry about meta-tags ever again -- our backend will add all missing fields, compose JSON-LD, and OG link-preview image on the fly.

Shipped with everything what you'd expect from mature network provider:
- Global scale CDN and caching for static assets
- WebSec: DDoS, WAF, OWASP
- Instant DNS
- Edge computing (workers)
- Analytics and insights for human users and AI agents
- Teamwork
- And of course "one click" migration form Cloudflare

Join /r/BridgeCDN and our waitlist - bridge-cdn.com

u/dr-dimitru — 7 days ago