← Articles

We scanned 399 local businesses. Only 7% got named by the AI we asked.

13 min read

Most of what is written about AI visibility for local business is either a vendor’s marketing or a national-brand study being stretched to cover a dental practice in Charlotte. We went and measured it instead: 399 local businesses across two North Carolina metros, 43 of them put through a full check of what Perplexity says when a customer asks for a recommendation, and then ten questions asked of three different assistants to find out how much of the answer depends on which one you ask.

Some of it supports the pitch you have probably heard. Some of it does not, including one finding that undercuts a fix we sell and another that undercuts the way our own free check tends to be read. All of it is below with the method and the sample sizes attached, so you can judge how much weight any of it carries.

This is a findings piece and it assumes the premise. If you are still working out what the category is, what GEO means for a local business, and whether yours needs it is the primer, and the better place to start: a fair number of businesses read these numbers and correctly conclude the work can wait.


What we measured, and how

Three datasets, gathered between August 25 and September 10, 2026.

The site scan (n=399). Each business came out of a Google Places search for its city and trade, filtered to those with at least 25 reviews. We fetched the homepage, read its JSON-LD structured data, read its robots.txt, and then requested the homepage again once per AI crawler, sending that crawler’s User-Agent, to see which ones the server answers. 349 of the 399 answered an ordinary browser request, and 386 completed the crawler probe, the larger number because a site that refuses a crawler has still given us a result.

The Perplexity answers (n=43). For 43 Charlotte dental practices we ran the real question a customer would ask, “best dentist in Charlotte, NC” and close variants, and kept the full answer, every source it cited, and every business it named.

The three-engine corpus (n=30). Ten Charlotte dental questions, from broad (“best dentist in Charlotte, NC”) to specific (“best Invisalign in South End, Charlotte NC”, “dentist in Charlotte, NC that takes Delta Dental”), each asked once of Perplexity, of OpenAI’s web-search API, and of Gemini with Google Search grounding. Thirty answers, 395 citations, no failed calls.

Seenvia measurement. Site scan: 399 businesses across Charlotte NC dental (220, checked 2026-08-31), Raleigh NC dental (145, 2026-09-01) and Charlotte NC plumbing and HVAC (34, 2026-09-02). Perplexity answers: 43 Charlotte dental practices, 2026-08-25 to 2026-08-31. Three-engine corpus: 2026-09-10. Google-Extended and Applebot-Extended are read from robots.txt only and left out of the live-request figures, because Google and Apple both state they have no HTTP user agent to send.


Finding 1: 93% are invisible, and the cast barely changes

Of the 43 practices we ran a real Perplexity answer for, 3 were named, or 7%. Two were cited as a source, 5%. The other 93% did not appear in the answer their own customers would see.

On its own that is not much of a finding, since you would expect a short answer to name a handful of businesses out of hundreds. The interesting part is how little the short list moves. We asked the identical question, “best dentist in Charlotte, NC”, 35 separate times. Across the 595 pairs of runs that allows, the sets of cited sources overlapped by an average of 0.78 on a scale where 1.00 means identical every time. 25 distinct sources were cited in total, and 8 of them turned up in every run.

Practice names move around more than sources do, and the two are worth separating. Only 18% of run pairs gave back the identical set of names, and only 6% gave the same names in the same order. That is the same instability SparkToro and Gumshoe.ai measured at national-brand scale, though a milder version of it: they found under a 1 in 100 chance of an identical brand list across two runs, against roughly 1 in 6 here. Underneath the churn the cast is close to fixed. Charlotte Dentistry, Piedmont Dentistry and Ballantyne Dentistry each appeared in 34 of the 35 runs.

Across all 43 answers, 37 practices were named at least once out of 219 total name-mentions, and the top five took 74% of them: Ballantyne Dentistry (38), Charlotte Dentistry (35), Piedmont Dentistry (34), East Blvd Dentistry (28) and SouthPark Dental and Oral Care (27).

That takes away the most comforting story about AI answers, the one where they are random and today’s absence is tomorrow’s appearance. The order a practice appears in really does shuffle. Whether it appears at all does not, and for almost everyone the answer is no.


Finding 2: Perplexity is reading other people’s pages about you

Across the 43 answers, Perplexity cited 875 sources from 81 distinct hosts, a median of 21 sources per answer. We went through all 81 by hand and sorted them on one question: does this page serve one business, or many?

54% of citations went to a page that serves many. Directories, roundup lists, insurance find-a-dentist pages, a local magazine, Reddit. In the median answer the share was 60%.

Cited sourceCitation slotsWhat it is
healthgrades.com101Provider directory
reddit.com46Four r/Charlotte threads, two of which carry 43 of the 46
charlottedentistry.com41A practice website
yellowpages.com41“Best 30 Dentists in Charlotte, NC”
ballantynedentistry.com41A practice website
deltadental.com36Insurance find-a-dentist
piedmontdentistrysmiles.com36A practice website
zocdoc.com36Booking directory
opencare.com36“20 Best Dentists Near Me in Charlotte”
doctor.webmd.com35“Best Dentists in Charlotte, NC (2026)”
toothcompass.com35“Best Dentists in Charlotte, NC (2026)”
charlottemagazine.com33Local magazine “Find a Dentist”

Seenvia measurement, 875 citation slots across 43 Perplexity answers, Charlotte dental, 2026-08-25 to 2026-08-31. A slot is one source in one answer, so a host cited twice in an answer counts twice: healthgrades.com appeared in 36 of the 43 answers and took 101 slots.

Four of those twelve rows are practice websites, and they stay in the table because taking them out would make it argue something the data does not. The top of the list is mixed. Most of it is still somebody else’s best-dentists-in-Charlotte page, with one directory well clear of everything else, but practice sites do hold their own inside it.

What the assistant is not doing is crawling a couple of hundred dental websites and forming a view of its own. It reads the roundups, the directories and a few Reddit threads, then summarises who those sources already agree on, which is why the practices that did get cited were mostly ones already sitting on those lists. The chain does not run from your website to the assistant. It runs from your reputation, to the lists and directories that rank on it, to the assistant, with your website off to one side of it.


Finding 3: the three assistants barely agree with each other

Everything above is one assistant. Before publishing any of it we wanted to know how much survived contact with the others, so we put ten Charlotte dental questions to Perplexity, to OpenAI’s web-search API and to Gemini’s Google Search grounding, and compared the sources each one cited for the same question.

They almost never matched.

ComparedMean overlap of cited domains
Perplexity vs Perplexity, same question re-asked0.78
Perplexity vs Gemini0.24
Perplexity vs ChatGPT (OpenAI web search)0.12
ChatGPT vs Gemini0.10

Seenvia measurement. Jaccard overlap of the set of cited domains, 1.00 = identical. Cross-engine rows: 10 questions, one run per engine, 2026-09-10. The first row is the repeat-to-repeat baseline from Finding 1, 35 runs of one question on Perplexity between 2026-08-25 and 2026-08-31, and it is what makes the other three readable: two engines overlap three to eight times less than two runs of the same engine do.

127 distinct domains were cited across the three engines and 13 of them were cited by all three. They do not even cite the same number of sources: Perplexity a median of 20 per answer, Gemini 16, OpenAI’s web search 5.

Finding 2 does not transfer either. Running the same hand classification over this corpus, the share of citations going to a page that serves many businesses was 37% for Perplexity and 32% for Gemini, against 0 of OpenAI’s 47 citations. On these ten questions it cited practice websites and almost nothing else. Narrow it to the two broadest questions, the ones closest to what the 43-answer study asked, and the split is 66% for Perplexity, 45% for Gemini, still 0% for OpenAI.

The one place they converge is at the very top, where a single practice, Ballantyne Dentistry, sits in all three engines’ five most-cited practice websites. Below that the lists have almost nothing in common.


Finding 4: what actually predicted being cited

Having both the scan and the answers let us ask the question the industry mostly asserts an answer to: did the businesses with the better technical setup actually get cited more?

We took the 195 reachable Charlotte dental practices, split them by whether Perplexity cited their domain in at least one of the 43 answers (39 cited, 156 never), and compared them.

SignalCited (n=39)Never cited (n=156)p
Has JSON-LD structured data84.6%82.1%0.71
Has LocalBusiness schema64.1%53.8%0.25
Blocks an answer-path AI crawler0.0%6.4%0.11
Review count (median)4913380.019

Seenvia measurement. Two-proportion z-tests on the three binary signals; tie-corrected Mann-Whitney U on review counts. Charlotte dental, reachable sites only. “Cited” means the practice’s own domain appeared as a source in at least one of the 43 answers.

The first two rows are the ones we did not expect. Having structured data at all showed essentially no relationship to being cited (p=0.71). The more specific LocalBusiness schema was about 10 points better among cited practices, which points the right way but is nowhere near significant at this sample size (p=0.25).

We sell a fix for LocalBusiness schema. Our own data does not currently support leading with it, so we have stopped leading with it.

One signal separated the groups. Cited practices had a median of 491 reviews against 338 for the never-cited (p=0.019). That is the only clearly significant difference we found, and it fits Finding 2: reviews are what the directories and roundup lists rank on, and those are what the assistant reads.

A second pointed the right way without getting there. None of the 39 cited practices blocked a crawler on the AI answer path, against 6.4% of the never-cited group, but zero in one group is exactly the shape a sample this size cannot resolve and p=0.11 reflects that. Treat it as a reason to check your own site rather than as a mechanism we have demonstrated. It is also the only item here you can fix this afternoon.


Finding 5: most “AI is blocked from your site” warnings are false alarms

Run an AI-readiness checker over a local business and there is a good chance it reports that AI crawlers are being blocked. In our probe of 386 sites, 48% turned at least one AI crawler away. That is an alarming number, and it is also close to useless.

AI crawlers do not all do the same job. Some fetch pages to build the index an assistant answers from, or open a page because a person has just asked something that needs it: OAI-SearchBot and ChatGPT-User for ChatGPT, PerplexityBot and Perplexity-User for Perplexity, Claude-SearchBot and Claude-User for Claude. Others collect pages into training or research corpora and play no part in answering anything, among them GPTBot, ClaudeBot, CCBot, and Bytespider, ByteDance’s scraper.

Split the 48% along that line and it comes apart:

What the server turned awayShare of 386 probed sites
Any AI crawler at all48%
A crawler on the answer path13%
Only training / dataset crawlers (mostly Bytespider, CCBot)35%

Seenvia measurement, live User-Agent requests against the 386 sites that completed the probe, 2026-08-31 to 2026-09-02. A refusal means the server answered and turned the request down, usually a 403 but sometimes a 503 or a bot challenge page. Crawler roles are the ones in our own crawler registry, taken from each operator’s published documentation. Robots.txt rules are counted separately, below.

So 35% of these sites would be flagged over crawlers that have no bearing on whether an assistant can answer a question about them. Blocking Bytespider is a normal, deliberate choice, and one many hosts and CDNs make by default. Counting it as an AI-visibility problem makes the problem look more than three times bigger than it is.

The same trap sits inside Cloudflare’s managed robots.txt, which disallows the training crawlers by name (GPTBot, ClaudeBot, CCBot, Bytespider, Google-Extended, Applebot-Extended) while leaving everything else under a blanket Allow, and states the intent outright with a search=yes, ai-train=no content signal. A checker that counts robots.txt entries without reading what they are for will tell that site’s owner that AI cannot see them. It can.

Count only the crawlers on the answer path, across both robots.txt rules and live server behaviour, and 15% of the 386 had a real AI access problem.

One judgement call sits inside that number. Google-Extended is a robots.txt token with no user agent of its own, and it governs Gemini grounding and model training together. We do not count it on the answer path, here or in the product, because we are not willing to tell a small business to opt into model training on evidence we do not have. Count its robots.txt rule and the figure is 16% rather than 15%.


Limits — read these before you quote the numbers

We would rather you cite this accurately than widely.

  1. One assistant carries most of the weight. Findings 1, 2 and 4 all rest on Perplexity. Finding 3 is the reason that matters, and it is a small study in its own right: ten questions, one run per engine.
  2. One market, one industry, for the answer data. Every answer we collected is Charlotte, NC dental. The 399-site scan spans two metros and two trades. The correlations do not.
  3. n=195 for the correlations. With 39 cited businesses, this study can detect large effects and would miss moderate ones. A null result here is weak evidence, not proof of no effect.
  4. “Cited” is a proxy. It means the practice’s domain appeared among roughly 21 sources for one question. It is not the same as being recommended, and not the same as getting a patient.
  5. The crawler probe is not the crawler. Our requests come from our own servers, not from OpenAI’s or Anthropic’s IP ranges, and Cloudflare and its competitors verify real crawlers by signed request, published range or reverse DNS. A refusal we recorded is consistent with a rule aimed at that crawler and with a rule aimed at anything claiming to be it.
  6. Sorting sources is a judgement. We labelled all 81 cited hosts in Finding 2, and all 127 in Finding 3, by hand: does this page serve one business or many? Reasonable people would move a handful of the edge cases, which is why the rule is stated rather than assumed.
  7. Correlation, not causation. Nothing here was an experiment. Practices with more reviews differ from practices with fewer in plenty of ways besides review count.
  8. A snapshot. Late August and early September 2026. These systems change monthly.

What we would actually do with this

Ranked by what the data supports, not by what is easiest to sell:

  1. Find out whether you are on the lists AI reads. Ask the question your customers would ask, read the sources cited rather than the answer itself, and check whether you appear on those specific pages. For Charlotte dental on Perplexity that meant Healthgrades, a WebMD roundup, Opencare, a local magazine and a couple of Reddit threads. Your market’s list will be different, which is the reason to read your own sources instead of borrowing ours. The step-by-step version of that, including how to read the sources panel, takes about thirty minutes.
  2. Do it in more than one assistant. This is the practical half of Finding 3. The sources behind a ChatGPT answer and a Perplexity answer to the same question had almost nothing in common, so checking one and generalising is how people end up confidently wrong about where they stand.
  3. Keep earning reviews. Slow, and the only thing in our data with a clearly significant relationship to being cited. It is also what the roundup lists rank on, so it compounds through the same channel.
  4. Check whether you are blocking the crawlers that matter, meaning the ones on the answer path rather than the training scrapers. This applies to a minority of sites, but where it applies it is real and free to fix. Our AI crawler checker runs the same two-layer check Finding 5 is built on: the robots.txt rules, then one live request per crawler.
  5. Fix your structured data, but stop expecting it to be the lever. It is cheap, it is correct practice, and it may well matter more than we could detect at this sample size. It is not where we would spend first.

The uncomfortable implication of Finding 2 is that a good part of local AI visibility is not a website problem at all. It is a reputation-and-listings problem, and it looks a lot more like the unglamorous work of being well regarded in your city than like anything you can install.