Same brand, same question, different country = different AI answer. And switching the language of the prompt changed it again. (what we're seeing tracking location-based AI visibility)

Same brand, same question, different country = different AI answer. And switching the language of the prompt changed it again. (what we're seeing tracking location-based AI visibility)

https://preview.redd.it/o0pkbp47gdkh1.png?width=2752&format=png&auto=webp&s=02225aa1080fb3a291796ebb6b6f38d2c5fe2d3b

Been tracking AI answer visibility per-location for clients and wanted to share a pattern that keeps surprising people, because it breaks the assumption that "our AI visibility" is one thing.

The setup: same brand, same buyer-intent prompt, run against the same engine, but varying (a) the location signal and (b) the language of the prompt. Two findings that changed how we think about this:

1. The answer changes by location, not just the ranking, the whole cited-source set.
Ask an engine a category recommendation question as a user in one country vs another and you don't just get a reordered list, you get different brands surfaced and a different set of sources cited to justify them. The model is filtering its retrieval by the location it infers, so each region is effectively drawing from its own corpus of reviews, listings, and local pages. A brand that's the confident #1 answer in one market can be absent in another off the identical prompt. For any multi-location or multi-market brand that means a single national/global "visibility score" is basically meaningless, you're averaging over answers that don't resemble each other.

2. Language of the prompt is a separate variable from location, and it moves the answer independently.
This one caught us off guard. A client operating in Finland: we ran the category prompts in English, then ran the same intent in Finnish. Different answers. Not just translated, different brands cited and different sources pulled. Our read is that the Finnish-language query pulls from a different slice of the corpus (Finnish-language reviews, local pages, local forum/press content) than the English version of the "same" question does, even for a user in the same place. So "location" and "prompt language" are two separate levers, and if you only ever test in English you're blind to what your actual local-language buyers are seeing.

The practical takeaway: if you operate in more than one country or more than one language, you have to measure per-location and per-language, on native-language prompts, not a translated English set. The gaps show up in exactly the places an English-only audit can't see.

Disclosure per rule 5: this comes out of our own tool (sanbi.ai, we do per-location AI visibility tracking), so that's where the data's from, weigh it accordingly. But you can sanity-check the effect yourself for free, ask ChatGPT or Perplexity a category question with different location context, then ask it once in English and once in the local language, and watch the cited sources change.

Curious if others tracking this see the language effect too, or whether it's stronger in some languages than others. My hunch is it's biggest in markets with a rich native-language web (Finnish, Japanese, German) and smaller where the local audience mostly consumes English content, but I only have a handful of markets to go on.

reddit.com
u/Sanbi_Ai — 13 hours ago
▲ 1 r/aeo

Same brand, same question, different country = different AI answer. And switching the language of the prompt changed it again. (what we're seeing tracking location-based AI visibility)

https://preview.redd.it/f7kh1mz58dkh1.png?width=2752&format=png&auto=webp&s=aed3551687b1cebc477e1f2f78b794b8e5513aca

Been tracking AI answer visibility per-location for clients and wanted to share a pattern that keeps surprising people, because it breaks the assumption that "our AI visibility" is one thing.

The setup: same brand, same buyer-intent prompt, run against the same engine, but varying (a) the location signal and (b) the language of the prompt. Two findings that changed how we think about this:

1. The answer changes by location, not just the ranking, the whole cited-source set.
Ask an engine a category recommendation question as a user in one country vs another and you don't just get a reordered list, you get different brands surfaced and a different set of sources cited to justify them. The model is filtering its retrieval by the location it infers, so each region is effectively drawing from its own corpus of reviews, listings, and local pages. A brand that's the confident #1 answer in one market can be absent in another off the identical prompt. For any multi-location or multi-market brand that means a single national/global "visibility score" is basically meaningless, you're averaging over answers that don't resemble each other.

2. Language of the prompt is a separate variable from location, and it moves the answer independently.
This one caught us off guard. A client operating in Finland: we ran the category prompts in English, then ran the same intent in Finnish. Different answers. Not just translated, different brands cited and different sources pulled. Our read is that the Finnish-language query pulls from a different slice of the corpus (Finnish-language reviews, local pages, local forum/press content) than the English version of the "same" question does, even for a user in the same place. So "location" and "prompt language" are two separate levers, and if you only ever test in English you're blind to what your actual local-language buyers are seeing.

The practical takeaway: if you operate in more than one country or more than one language, you have to measure per-location and per-language, on native-language prompts, not a translated English set. The gaps show up in exactly the places an English-only audit can't see.

Disclosure per rule 5: this comes out of our own tool (sanbi.ai , we do per-location AI visibility tracking), so that's where the data's from, weigh it accordingly. But you can sanity-check the effect yourself for free, ask ChatGPT or Perplexity a category question with different location context, then ask it once in English and once in the local language, and watch the cited sources change.

Curious if others tracking this see the language effect too, or whether it's stronger in some languages than others. My hunch is it's biggest in markets with a rich native-language web (Finnish, Japanese, German) and smaller where the local audience mostly consumes English content, but I only have a handful of markets to go on.

reddit.com
u/Sanbi_Ai — 14 hours ago

A cheap due-diligence signal I've started checking on retail/QSR tenants: their per-location "AI answer" visibility

not sure if this belongs here or if I'm overthinking it, so tell me if it's noise. it's about underwriting and monitoring franchise / QSR / service tenants, so figured it's at least CRE-adjacent.

context: a growing share of "where should I eat / who should I call / is this place any good" now happens inside AI answers (ChatGPT, Google's AI Overviews, Perplexity, Siri) instead of ten blue links. for the local retail and service tenants a lot of us underwrite, that's a real slice of top-of-funnel discovery. and here's the thing that surprised me: it varies wildly per unit, even inside the same brand.

why it varies is kind of mundane. different engines pull local data from different places:

  • ChatGPT leans on Yelp + Bing local
  • Google's AI Overviews pull from Google Business Profile + Maps reviews
  • Siri routes through ChatGPT / Apple Maps
  • Perplexity crawls the web and cites sources

so two units of the same franchise can be described completely differently depending on which of those listings is claimed, current, and well-reviewed at that specific address. the brand can be strong nationally and a specific box can still be functionally invisible or described badly.

where I think this actually matters for us:

  1. tenant health / sales durability. for a percentage-rent deal or any tenant where you care whether the unit keeps its sales, per-location review sentiment and AI-answer visibility is a leading indicator that's cheaper to check than a sales audit. a unit whose local discovery is quietly degrading is a unit whose foot traffic may follow.
  2. underwriting a franchise tenant. everyone checks the franchisor's financials and the guarantor. fewer people check whether THAT specific location has a claimed, healthy local footprint, or whether it's a dead Yelp/GBP listing that AI models skip. costs ten minutes and tells you something about the operator.
  3. brand-level contagion risk if you hold several units of one brand. AI models don't cleanly separate one bad location's reviews from the brand. a few chronically bad units in a system can drag how AI answers generic "is [brand] worth it" questions, which theoretically softens demand across the brand you're exposed to. this one I'm least sure about, would push back on myself here.

the check itself is dumb-simple and free: ask ChatGPT and pull up Google's AI Overview for "best [category] near [address/neighborhood]" and "is [brand] at [location] good," and look at whether your tenant shows up, and how it's described. then glance at the GBP/Yelp review recency for that unit. degrading review recency + absence from the AI answer is at least a yellow flag worth a question to the operator.

honest disclosure: this overlaps with my day job (I work at a company that tracks this stuff per location), so I'm obviously biased toward thinking it matters. which is exactly why I'm posting here instead of assuming, this sub is the right crowd to tell me if it's real signal or if I've got a hammer looking for nails.

so: does anyone actually factor local-discovery health into tenant underwriting or asset monitoring yet? or is it too early / too soft to bother, and NOI and guarantor strength are still the only things that move the needle? genuinely want the pushback.

reddit.com
u/Sanbi_Ai — 1 day ago

AI Search (ChatGPT, Gemini) gives completely different answers depending on your city. Here is why this matters for local Small Businesses.

For the last year, everyone tracking AI visibility has been asking: "Does ChatGPT mention my business?"

That is the wrong question.

We ran thousands of identical prompts across ChatGPT, Gemini, Perplexity, and Claude from different geographic contexts. The results confirmed that AI answers are not the same in every city. Across the prompts we tested, the top-recommended product or service changed in 41% of major U.S. metros for the exact same query.

For categories like home services, fitness, and local retail, the variance was even higher.

If you are running a small business, this is a critical shift. When a user asks an AI assistant for a recommendation, the model does not pull from a single global ranking. It blends:

  1. Localized retrieval (Google and Bing SERPs return different local packs by region)
  2. Regional citation sources (local publications, local reviews, city-specific forums)
  3. Inferred location signals (user IP, prompt context like "near me")

The Google Business Profile (GBP) Angle

This means your Google Business Profile and local citations feed directly into the AI's localized logic. A business that dominates the AI response in one ZIP code can be completely invisible just a few miles away. We call this "regional drift."

If your small business relies on local foot traffic or service areas, you cannot rely on a generic, national AI visibility score. You are flying blind. The AI is heavily weighing where you are, using your GBP data and local directory mentions to filter you in or out of the response.

We just launched a tool (Sanbi AI) to map this out geographically, allowing brands to see their AI visibility as a literal map instead of a single score. But regardless of the tools you use, the takeaway for small businesses is clear: localized content, geo-targeted reviews, and consistent GBP signals are what dictate if an AI recommends you to a local buyer.

Has anyone else noticed their business showing up inconsistently in AI responses depending on where the prompt is run?

u/Sanbi_Ai — 3 months ago

Founders use AI for everything, but still pay agencies $4k a month for AEO. I built a tool to replace them in 30 minutes a week.

A year ago, I noticed a weird disconnect. Small business owners and startup founders were using Claude to write proposals and ChatGPT to handle customer support. They were completely self sufficient on hard knowledge work. But they were still paying SEO and AEO agencies $3,000 to $5,000 a month just to get a dashboard and a monthly report on their AI visibility.

If you have AI tools, you do not need an agency for this. You just need the right data to point your efforts at.

I decided to build a platform that makes AEO affordable for SMBs who do not have enterprise budgets. It is called Sanbi AI, and the entry plan is $50 a month.

How it works

The core idea is to automate the tracking so you only spend 30 minutes a week actually moving the needle.

  1. Automated Monitoring: The system runs your buyers' exact intent queries (like "best CRM for realtors") into ChatGPT, Claude, Gemini, and Perplexity on a strict schedule.
  2. Source Extraction: It finds exactly which domains the AI engines are citing to build their answers.
  3. The Growth Pipeline: It cross references where your competitors got cited and you did not, dropping these missed opportunities into a pipeline.
  4. AI Drafting: You click a missed citation, and the tool drafts a context aware reply or piece of content to help you get cited on that exact source.

What I learned building this

Tracking AI visibility at scale breaks a lot of classic SEO rules. Here are the biggest surprises from building the platform:

  • Reddit is the most important domain: Across the engines we track, Reddit threads show up in AI answers significantly more often than brand websites. If you are a SaaS or service business and you are not in relevant subreddit discussions, you are practically invisible to AI search.
  • Perplexity has a massive YouTube bias: A huge chunk of Perplexity citations are YouTube videos, often only tangentially relevant. We actually had to build a custom filter to hide these "cheap citations" because they were drowning out high value blog and article mentions.
  • Prompt variance is brutal: If you ask Claude the exact same prompt five times in an hour, you will get different brand lineups. Manually checking your visibility once is useless. You have to sample it on a schedule over time to see your actual share of recommendations.

The results so far

We currently have about 40 paying customers working the pipeline. On average, early users are finding 15 to 30 concrete citation opportunities in their first week. These are specific URLs where AI is recommending a competitor and the user has a realistic shot at inserting their brand into the conversation.

I am happy to answer any questions about how AI engines decide who to cite, how we built the scraping infrastructure, or if AEO makes sense for your specific category. (Also we have a free tier!!)

- Sanbi.ai

u/Sanbi_Ai — 3 months ago

I built a pipeline to track AI search visibility across 4 engines and automate AEO. Here is how it replaces a $4k agency retainer.

Most of us here know that standard SEO is bleeding traffic to AI engines like Perplexity, Claude, ChatGPT, and Gemini. If a user asks "best CRM for startups" and your tool or your client is not in the output, you do not exist to that buyer.

The problem is that tracking this manually is impossible because prompt outputs change every hour. Answer Engine Optimization (AEO) agencies are currently charging $3k to $5k a month just to run these checks and write basic content.

I wanted to automate this entire loop for myself and other founders. So I built Sanbi AI. The goal was to make a tool where a single operator could manage their entire AI visibility pipeline in 30 minutes a week, starting at $50 a month.

Here is the architecture of the Growth module we just launched:

  • Prompt Monitoring: The system runs buying intent queries into the 4 major AI engines on a strict schedule to account for response variance.
  • Source Extraction: It pulls every source URL the AI cited to build its answer (blogs, Reddit threads, news).
  • Citation Pipeline: It cross references where your competitors got cited and you did not, dropping these missed opportunities into a pipeline board.
  • Automated Drafts: You click a missed citation, and the tool uses AI to draft a context aware reply or post to help you get cited on that exact source domain.

Everything flows through an SSR proxy, and scores are computed live based on citation count, recency, and platform authority.

The core thesis is that operators armed with the right automation can bypass expensive agencies entirely.

I am curious how the builders and marketers here are handling AI search visibility right now. Are you building your own scrapers, ignoring it until the dust settles, or paying for enterprise tracking?

- Sanbi AI

reddit.com
u/Sanbi_Ai — 3 months ago

Now that founders are using ChatGPT and Claude to do their own copywriting, research, and strategy, is the $3 to $5k/mo agency model dead for SMBs? Or am I overplaying this?

Context: I'm building Sanbi AI, an AEO and AI search visibility platform (flair discloses). The reason I started it: a year ago I watched a small agency owner pay an SEO/AEO consultancy $4k/mo for what was basically a dashboard, a monthly report, and some content drafts. Same month, that same person was using Claude to write proposals, ChatGPT to draft client emails, and Perplexity to research competitors. The disconnect was wild. They were happily self sufficient on hard knowledge work but outsourcing the part that's mostly tracking and templated output.

So I built Sanbi AI around the assumption that an SMB owner with AI tools and roughly 30 minutes a week can do the work an AEO agency charges four figures a month for, if the tool surfaces the right opportunities and drafts the first pass. Entry plan is $50/mo against an agency rate of $3 to $5k. Just launched the Growth module, which pipelines citation opportunities (the URLs where AI engines are recommending your competitor instead of you across ChatGPT, Claude, Gemini, and Perplexity), drafts the outreach and content, and lets the owner work the pipeline themselves. Early beta feedback is encouraging but it's still beta.

The strategic question I keep going back and forth on:

Across marketing categories generally (SEO, AEO, content, PPC management, social), how much of the agency retainer market do you actually think gets eaten by "owner plus AI tool plus 30 min/week" in the next 2 to 3 years?

Three positions I've heard from people I respect:

  1. Most of it. Agencies were always selling time and templated outputs. AI collapses both. The retainer model below roughly $10k/mo basically goes away for anything that isn't deep strategy or relationship based.
  2. Almost none of it. Owners say they want to DIY but won't actually do 30 min/week consistently. Agencies survive because they're a forcing function as much as a service. DIY tools have been around for years (Ahrefs, SEMrush, HubSpot) and agencies still grew.
  3. Splits by category and owner type. Technical and measurable work (rank tracking, audits, AEO monitoring, reporting) gets eaten fast. Creative and relationship work (brand strategy, PR, partnerships) doesn't. And it splits by founder. Operators who already use AI daily will DIY. Owners who don't use AI at all keep paying agencies.

I'm betting on position 3, which is why I priced Sanbi AI for self serve SMBs instead of building another enterprise dashboard. But I'd like a reality check from people who actually run agencies, sell to SMB owners, or have watched this play out in other categories.

For agency owners specifically: are you already seeing churn from clients saying "we're going to try this in house with AI for a while"? And for SMB owners: have you actually pulled work back in house, or is it still mostly aspirational?

Genuinely curious where I'm wrong on this.

reddit.com
u/Sanbi_Ai — 3 months ago

We tracked 42 B2B brands across ChatGPT, Claude, Gemini & Perplexity for 74 days — only 19% appeared consistently in all four. Cross-engine AEO is not one channel, it's four.

Disclosure up front: I'm the founder of Sanbi.ai, an AEO tracking platform. The data below is from our own monitoring infrastructure. Methodology and limitations are spelled out so you can poke holes in it.

TL;DR

Most AEO strategy decks treat "AI search" as a single channel. Our data says it isn't. Cross-engine citation overlap for the same brand on the same prompts is far lower than the industry talks about. If you're optimizing as if ChatGPT, Claude, Gemini, and Perplexity behave alike, you're leaving 60–80% of the surface uncovered.

What we looked at

  • 42 B2B brands across SaaS, professional services, and local service categories
  • ~280 unique buying-intent prompts (avg 6–8 per brand), e.g. "best [category] for [use case]", "alternatives to [competitor]", "top [service] in [city]"
  • 4 engines: ChatGPT, Claude, Gemini, Perplexity
  • 74-day window (Mar–May 2026), each prompt re-sampled multiple times per week to control for response variance
  • Definition of "cited": brand mentioned by name in the AI's answer body, not just in linked sources

Key finding: cross-engine overlap is low

Of the 42 brands tracked:

Cited in… % of brands
All 4 engines (consistent) 19%
3 of 4 engines 26%
2 of 4 engines 33%
Only 1 engine 17%
0 engines (for their tracked prompts) 5%

Translation: ~55% of brands we tracked are essentially invisible in at least half the AI surface for queries their own buyers are asking. Most of them did not know this before they started measuring.

Where the gaps cluster

Looking at which engine each brand was missing from, two patterns kept repeating:

  1. Perplexity is the outlier. Brands well-cited in ChatGPT + Gemini + Claude were missing in Perplexity in 41% of cases. Perplexity weights recency and specific source types (YouTube, Reddit, news) much more heavily than the other three. A brand with strong evergreen SEO content but no recent press or community discussion tends to get skipped.
  2. Claude and ChatGPT overlap heavily (~71% co-citation rate when either cites the brand), but Gemini is more "Google-shaped" — i.e., it tracks Google SERP citations more closely than ChatGPT/Claude do. If you've been doing classical SEO well, Gemini is usually your strongest engine and Perplexity your weakest, by a wide margin.

Practical implication: AEO strategy needs per-engine plays, not a single content push.

  • For Perplexity: prioritize recency (publish new dated content monthly), Reddit thread presence, and being mentioned in news/PR within the last 90 days
  • For ChatGPT + Claude: prioritize source diversity — Reddit threads, third-party blogs, listicles, directory pages all matter more than your own domain
  • For Gemini: classical SEO + Google SERP presence is still doing most of the lifting

Second finding worth flagging: prompt variance is brutal

We re-ran the same prompt against the same engine multiple times within a 24-hour window. Brand mention sets changed in 38% of cases. Same query, same engine, different brands cited.

Implications:

  • Any AEO audit that runs a prompt once and reports findings is producing noise
  • You need ≥5 samples per prompt over time to get a stable signal on "share of recommendations"
  • This is also why "I asked ChatGPT and we showed up" anecdotes are basically useless as evidence

Third finding: Reddit dominates citations more than most decks acknowledge

Of the source URLs the engines cited when answering our 280 prompts, Reddit threads accounted for 22% of total citations — more than any other domain category, and more than 3x the share of the brands' own websites (~6.5%).

For B2B SaaS specifically, the share was even higher (~28% Reddit).

If your AEO content plan is heavy on owned-domain blog posts and light on community presence, you're optimizing for the smaller pie.

Limitations / what I'd want to test next

  • Sample size is 42 brands. Directionally interesting, not definitive. Would love to compare notes if anyone here has run similar tracking at higher volume.
  • Categories were skewed toward SaaS and professional services. Consumer/ecom would likely look different.
  • Engines update behavior monthly. The 41% Perplexity gap may compress over the next quarter as it normalizes.
  • We didn't separate "branded prompts" (where the brand is named in the query) from "category prompts." Including only category-level prompts would probably show even lower overlap.

Where this data came from

We pulled it from Sanbi.ai's monitoring infrastructure — the prompt sampling, engine breakdown, source domain aggregation, and citation tracking are core features of the product. If anyone wants to replicate the methodology on their own client base, the starter plan is $50/mo and the cross-engine breakdown + source domain analysis are in there. Link in profile, not dropping it inline since this sub explicitly asks for commentary not link-drops.

Mostly though, I posted this because the 19% / 41% / 22% numbers genuinely surprised me when I aggregated them, and I'd be curious whether anyone else tracking this stuff sees similar shapes or completely different ones. The cross-engine overlap question feels under-discussed relative to how much it should change strategy.

Happy to share more cuts of the data in comments if useful — by category, by engine pair, by source type, whatever's interesting.

u/Sanbi_Ai — 3 months ago

[For Hire] AI Search Visibility & AEO for B2B brands — track if ChatGPT, Claude, Gemini & Perplexity recommend you (plans from $50/mo)

What we do

Sanbi.ai tracks whether AI search engines (ChatGPT, Claude, Gemini, Perplexity) mention your brand when your buyers ask buying-intent questions like "best [your category] for [their use case]" — and gives you a clear pipeline to fix it when they don't.

If your customers are using AI to shortlist vendors before they ever visit your site (they are), this is the channel you're not measuring yet.

Who it's for

  • B2B SaaS founders and marketing leads
  • Agencies who want to add AEO as a service line
  • Consultancies, law firms, accounting firms, medical/dental practices
  • Local service businesses competing on "best [service] in [city]" queries
  • Anyone whose buyers are asking AI before they ask Google

What you actually get

  • Prompt monitoring — we run your buyers' real questions into 4 AI engines on a schedule and track who gets cited (you vs. competitors)
  • Competitor tracking — up to 20 competitors, monitored side-by-side
  • Geographic dominance — see visibility by city, region, or country
  • Source domain analysis — every page AI is citing in your category (Reddit threads, blogs, directories, YouTube, news)
  • Citation opportunity pipeline — specific URLs where competitors are cited and you aren't, with AI-drafted reply/comment/content suggestions to get you added
  • Site Health audits — SEO + AEO + ARO scoring on your pages with prioritized fixes
  • Sentiment analysis — how AI is positioning your brand on a 0–100 scale, with risks and selling points
  • News & media mention tracking
  • AI assistant + AI-generated blog drafts with tone control

About 30 minutes a week of your time to work the pipeline.

Pricing

  • Starter: $50/mo — built specifically for SMBs and startups priced out of the $3–5k/mo agency market
  • Pro & Enterprise tiers available for teams that need seat management, more monitors, more credits, and the full Growth pipeline

Why us vs. an agency

Most AEO agencies are reselling dashboards and charging $3–5k/mo. We built the dashboard and skipped the markup. You drive the strategy in 30 min/wk; we surface the opportunities and draft the work.

How to start

  • Site: https://sanbi.ai
  • Free to explore before paying
  • DM me here for a walkthrough or to talk about whether this fits your category — happy to look at your brand's current AI visibility on a quick call before you commit to anything
u/Sanbi_Ai — 3 months ago
▲ 13 r/Franchising+5 crossposts

AI Search (ChatGPT, Claude, Gemini) gives completely different answers depending on your city. Here is why this matters for local Small Businesses.

For the last year, everyone tracking AI visibility has been asking: "Does ChatGPT mention my business?"

That is the wrong question.

We ran thousands of identical prompts across ChatGPT, Gemini, Perplexity, and Claude from different geographic contexts. The results confirmed that AI answers are not the same in every city. Across the prompts we tested, the top-recommended product or service changed in 41% of major U.S. metros for the exact same query.

For categories like home services, fitness, and local retail, the variance was even higher.

If you are running a small business, this is a critical shift. When a user asks an AI assistant for a recommendation, the model does not pull from a single global ranking. It blends:

  1. Localized retrieval (Google and Bing SERPs return different local packs by region)
  2. Regional citation sources (local publications, local reviews, city-specific forums)
  3. Inferred location signals (user IP, prompt context like "near me")

The Google Business Profile (GBP) Angle

This means your Google Business Profile and local citations feed directly into the AI's localized logic. A business that dominates the AI response in one ZIP code can be completely invisible just a few miles away. We call this "regional drift."

If your small business relies on local foot traffic or service areas, you cannot rely on a generic, national AI visibility score. You are flying blind. The AI is heavily weighing where you are, using your GBP data and local directory mentions to filter you in or out of the response.

We just launched a tool (Sanbi AI) to map this out geographically, allowing brands to see their AI visibility as a literal map instead of a single score. But regardless of the tools you use, the takeaway for small businesses is clear: localized content, geo-targeted reviews, and consistent GBP signals are what dictate if an AI recommends you to a local buyer.

Has anyone else noticed their business showing up inconsistently in AI responses depending on where the prompt is run?

u/Sanbi_Ai — 1 day ago

AI hallucinated a Reddit URL ID and almost sent my client to a highly inappropriate sub. Build validation layers.

I was building a product that uses LLMs to draft responses for forums and blogs where clients are cited, with the goal of increasing visibility. I was testing the system with real client data.

The AI generated a response citing what looked like a perfectly normal Reddit thread about automotive sensors.

I clicked the generated link to verify it. Instead of a technical discussion, I was instantly redirected to a completely unrelated and highly inappropriate subreddit.

After a brief moment of panic, I investigated what actually happened. The AI correctly generated the base URL and the text slug for the automotive post. However, it completely hallucinated the unique string of characters (the ID) at the very end of the URL.

Because of how Reddit routing works, the platform ignored the text slug and redirected the browser based solely on that hallucinated ID. By pure chance, that random ID belonged to an active, very unprofessional post.

If this had gone to production and a client clicked that link, it would have been catastrophic for trust and business.

This made me realize that just checking for a 404 status code is not a viable validation strategy. A link can resolve successfully and still be entirely wrong.

To fix this, I built a specific verification layer that tests and classifies every URL before a human sees it.

The classifier sorts the links into specific buckets:

  • verified: The link exists and matches the intended context.
  • unverifiable: Blocked by paywalls or forum logins.
  • hallucinated: Dead IDs or 404s.
  • suspicious: The link works, but there is a redirect mismatch. This is the crucial category that catches the exact error I experienced.

TL;DR: An LLM hallucinated the unique ID of a Reddit URL, which caused a redirect to a bizarre and inappropriate page despite having a clean URL slug. If you are building AI tools that generate links, you must build robust verification systems that check for redirect mismatches, not just broken links.

reddit.com
u/Sanbi_Ai — 3 months ago

AI hallucinated a Reddit URL ID and almost sent my client to a highly inappropriate sub. Build validation layers.

I was building a product that uses LLMs to draft responses for forums and blogs where clients are cited, with the goal of increasing visibility. I was testing the system with real client data.

The AI generated a response citing what looked like a perfectly normal Reddit thread about automotive sensors.

I clicked the generated link to verify it. Instead of a technical discussion, I was instantly redirected to a completely unrelated and highly inappropriate subreddit.

After a brief moment of panic, I investigated what actually happened. The AI correctly generated the base URL and the text slug for the automotive post. However, it completely hallucinated the unique string of characters (the ID) at the very end of the URL.

Because of how Reddit routing works, the platform ignored the text slug and redirected the browser based solely on that hallucinated ID. By pure chance, that random ID belonged to an active, very unprofessional post.

If this had gone to production and a client clicked that link, it would have been catastrophic for trust and business.

This made me realize that just checking for a 404 status code is not a viable validation strategy. A link can resolve successfully and still be entirely wrong.

To fix this, I built a specific verification layer that tests and classifies every URL before a human sees it.

https://preview.redd.it/c79egro7155h1.png?width=716&format=png&auto=webp&s=d51a4c0b6a10693b149cc35f018692cd775336f3

As you can see in the attached image, the classifier sorts the links into specific buckets:

  • verified: The link exists and matches the intended context.
  • unverifiable: Blocked by paywalls or forum logins.
  • hallucinated: Dead IDs or 404s.
  • suspicious: The link works, but there is a redirect mismatch. This is the crucial category that catches the exact error I experienced.

TL;DR: An LLM hallucinated the unique ID of a Reddit URL, which caused a redirect to a bizarre and inappropriate page despite having a clean URL slug. If you are building AI tools that generate links, you must build robust verification systems that check for redirect mismatches, not just broken links.

reddit.com
u/Sanbi_Ai — 3 months ago
▲ 14 r/WTFisAI

AI hallucinated a Reddit URL ID and almost sent my client to an NSFW sub. Build validation layers.

I was building a product that uses LLMs to draft responses for forums and blogs where clients are cited, with the goal of increasing visibility. I was testing the system with real client data.

The AI generated a response citing what looked like a perfectly normal Reddit thread about automotive sensors.

I clicked the generated link to verify it. Instead of a technical discussion, I was instantly redirected to a hardcore NSFW subreddit.

After a brief moment of panic, I investigated what actually happened. The AI correctly generated the base URL and the text slug for the automotive post. However, it completely hallucinated the unique string of characters (the ID) at the very end of the URL.

Because of how Reddit routing works, the platform ignored the text slug and redirected the browser based solely on that hallucinated ID. By pure chance, that random ID belonged to an active NSFW post.

If this had gone to production and a client clicked that link, it would have been catastrophic for trust and business.

This made me realize that just checking for a 404 status code is not a viable validation strategy. A link can resolve successfully and still be entirely wrong.

To fix this, I built a specific verification layer that tests and classifies every URL before a human sees it.

classifier results

As you can see in the attached image, the classifier sorts the links into specific buckets:

  • verified: The link exists and matches the intended context.
  • unverifiable: Blocked by paywalls or forum logins.
  • hallucinated: Dead IDs or 404s.
  • suspicious: The link works, but there is a redirect mismatch. This is the crucial category that catches the NSFW error.

TL;DR: An LLM hallucinated the unique ID of a Reddit URL, which caused a redirect to an NSFW page despite having a clean URL slug. If you are building AI tools that generate links, you must build robust verification systems that check for redirect mismatches, not just broken links.

reddit.com
u/Sanbi_Ai — 3 months ago

Founders & GTM folks: What's working for LinkedIn cold outreach right now? (YC Template vs. Short & Casual)

Hey everyone,

I'm currently refining my LinkedIn cold outreach strategy and wanted to get a reality check from other founders and sales folks who are doing this daily.

There's a classic YC outreach template (see the attached image) that is often held up as the gold standard for cold messaging. It follows a very distinct, multi-paragraph formula:

  • The Hook: A highly personalized reference (e.g., referencing a recent blog post or company milestone).
  • The Value Prop & Social Proof: A quick intro, what the product does, and dropping notable customer names.
  • Team Credibility: Highlighting past big-tech or relevant industry experience.
  • The CTA: A direct ask for a demo, proposing a specific day, and dropping a Calendly link.

While this structure is obviously proven for cold email, I’m curious how well it translates to LinkedIn InMails or connection request DMs today.

My questions for the group:

  1. Do you still follow this YC-style structure? Is a multi-paragraph message with background and social proof too heavy for LinkedIn, or does the deep personalization still win out?
  2. Or do you keep it simple? Do you prefer a much shorter, casual approach—something that just starts with a simple "Hi John Doe, saw you're leading [Team]..." and asks a quick, low-friction question to start a conversation?
  3. The Calendly drop: Are you including your scheduling link in the very first cold message like in the template, or do you wait until they actually reply?

Would love to hear what styles, lengths, and CTAs are actually converting for you right now. Thanks!

https://preview.redd.it/z33kn1sjdy4h1.jpg?width=984&format=pjpg&auto=webp&s=a2c87de6a9f9bb2e08794a5048d11dc6229f1b5d

reddit.com
u/Sanbi_Ai — 3 months ago

Founders & GTM folks: What's working for LinkedIn cold outreach right now? (YC Template vs. Short & Casual)

Hey everyone,

I'm currently refining my LinkedIn cold outreach strategy and wanted to get a reality check from other founders and sales folks who are doing this daily.

There's a classic YC outreach template (see the attached image) that is often held up as the gold standard for cold messaging. It follows a very distinct, multi-paragraph formula:

  • The Hook: A highly personalized reference (e.g., referencing a recent blog post or company milestone).
  • The Value Prop & Social Proof: A quick intro, what the product does, and dropping notable customer names.
  • Team Credibility: Highlighting past big-tech or relevant industry experience.
  • The CTA: A direct ask for a demo, proposing a specific day, and dropping a Calendly link.

While this structure is obviously proven for cold email, I’m curious how well it translates to LinkedIn InMails or connection request DMs today.

My questions for the group:

  1. Do you still follow this YC-style structure? Is a multi-paragraph message with background and social proof too heavy for LinkedIn, or does the deep personalization still win out?
  2. Or do you keep it simple? Do you prefer a much shorter, casual approach—something that just starts with a simple "Hi John Doe, saw you're leading [Team]..." and asks a quick, low-friction question to start a conversation?
  3. The Calendly drop: Are you including your scheduling link in the very first cold message like in the template, or do you wait until they actually reply?

Would love to hear what styles, lengths, and CTAs are actually converting for you right now. Thanks!

https://preview.redd.it/z33kn1sjdy4h1.jpg?width=984&format=pjpg&auto=webp&s=a2c87de6a9f9bb2e08794a5048d11dc6229f1b5d

reddit.com
u/Sanbi_Ai — 3 months ago

These two Harvard B2B sales lectures by Kent Summers are absolute gold for early-stage founders.

Want to put this out openly because I see too many of us struggling with the reality of enterprise sales. If you're building in the B2B space, you need to watch these two sessions from Harvard Alumni Entrepreneurs featuring Kent Summers. The guy is just phenomenal at what he does.

He breaks down the actual tradecraft of B2B sales and dispels the myth that you should just go out and hire a professional salesperson to solve your revenue problems early on. He argues that as a founder, you have to be the first one selling so you intimately understand the blocking issues and the exact pain points your product solves. He also goes deep into managing a weighted pipeline and proactively purging the "slow no"—those prospects who take up all your time but never actually pull the trigger.

Here are the links:
Session 1:http://youtube.com/watch?v=OaNi0dntHfU
Session 2:https://youtu.be/A8Pl17h7h2Q?si=muSrqgh_5x-ZR3R0

As a technical founder currently building an AI visibility audit platform, this material was a massive reality check. It actually reinforced exactly why my top priority right now is finding a co-founder with a strong B2B sales and marketing background. The mechanics of enterprise sales are a completely different beast from building the tech, and having a partner who intrinsically understands pipeline discipline and customer acquisition costs is make-or-break.

Highly recommend blocking out some time to watch these if you're trying to figure out your go-to-market motion!

u/Sanbi_Ai — 3 months ago
▲ 5 r/Business_in_China+1 crossposts

I audited 50 Western brands in Chinese AI. 47 didn't exist!! 🚨

For the folks who don't have time, here's why:
Google doesn't exist there. Neither does Reddit. ChatGPT is blocked. Gemini is blocked. Your backlinks don't count. Your press releases don't count. Your PDFs don't count. The crawlers reading your site have names you've never heard of. And the Chinese models answer from training data, not live search, so the content you publish today might not show up for months.
Different internet. Different rules. Different playbook.

Want me to get a bit more technical? Here's the full breakdown.
We ran a brand visibility audit across 6 Chinese tech hubs (Shanghai, Beijing, Shenzhen, Hangzhou, Guangzhou) for a US B2B client. Same prompts, two sets of models: Western LLMs vs. Qwen (Alibaba) and DeepSeek.
The results weren't just different in degree. They were structurally different.

  1. Different training corpus, different internet. Qwen and DeepSeek weren't trained on Reddit, Hacker News, or English trade press. They learned from Baidu-indexed content, Zhihu technical columns, eet-china -com, 21ic-com, elecfans-com.

If you're not on those platforms, you don't exist in Chinese AI answers. Your Google ranking is irrelevant.

  1. First-party content dominates. Western AEO leans on third-party authority like backlinks, reviews, analyst coverage. In Chinese LLMs, citation analysis tells a different story: your own domain gets cited more than any single trade publication.

Every product page with full specs in crawlable HTML. Every application note. Every design brief. That's the lever.

PDFs don't move the needle. Press releases don't either.

  1. Crawlers you've never heard of are already on your site. Baiduspider. QwenBot. DeepSeekBot. ByteDance crawlers.
    Two questions: Is your CDN blocking them? (Many sit on Alibaba Cloud and Tencent ASNs that default firewalls flag.) And are you serving zh-CN localized pages?

A missing hreflang tag is the single most common reason Western brands underindex in Chinese AI.

  1. Training cycles, not publishing cycles. Perplexity and Gemini do live web search. You can influence them this week. Qwen and DeepSeek answer mostly from training data.

That means your window to shape outputs is the next training cut, not your editorial calendar. Syndicate to the platforms they index now.

  1. ChatGPT is blocked. Gemini is blocked. Measuring "AI visibility" with only Western models for a China strategy is measuring the wrong thing.
    If you sell B2B into China, your AEO strategy needs a separate Chinese LLM track.

It's not a translation problem. It's a different internet.

reddit.com
u/Sanbi_Ai — 3 months ago