aiDeep AI Solutions Research

AI Research Division · AI Visibility

The 1.2% Problem

SOCi found ChatGPT recommends 1.2% of local business locations against 35.9% in Google's local 3-pack. The Houston AI Visibility Index measured why: across 750 prompts and 3,750 responses, machine-readable identity out-predicted backlinks by roughly 3.4x. Includes the Five-Engine Audit — a reproducible field protocol published without restriction.

Shayne Beavan

By Shayne Beavan

Founder, Deep AI Solutions · Inventor of record, 6 USPTO filings

14 min

There is a number in the 2026 local search data that should have caused more alarm than it did.

SOCi's 2026 Local Visibility Index evaluated more than 350,000 business locations across 2,751 multi-location brands, scoring each against 120+ metrics on six platforms. When they asked ChatGPT for local recommendations, it named 1.2% of those locations. The same businesses appeared in Google's local 3-pack 35.9% of the time.

Gemini recommended 11%. Perplexity, 7.4%. Eighty-three percent of restaurants did not appear in AI-generated local recommendations at all.

Put plainly: getting recommended by ChatGPT is roughly thirty times harder than earning prominent placement in traditional local search. And those were the large businesses. SOCi's methodology restricted the study to brands operating fifty or more locations — companies with marketing departments, budgets, and in many cases agencies already working the problem.

Meanwhile, BrightLocal's 2026 consumer research found that 45% of consumers now use AI tools to find local services. One year earlier that figure was 6%.

Demand moved roughly sevenfold in twelve months. Supply-side visibility sits near one percent. That is not a marketing gap. That is a market failing to clear.

The obvious question is why. We spent a month measuring it, and the answer turned out to be almost the opposite of what the industry is selling.

What we measured

The SOCi-class studies, excellent as they are, systematically exclude the businesses that constitute the actual local economy.

There are roughly 33 million small businesses in the United States. SOCi's universe was 2,751 brands. The single-location dentist, the owner-operated HVAC contractor, the three-truck roofing company sit in a measurement desert — nobody publishes rigorous data on them, because at that scale there is no existing dashboard to pull from.

So we measured them directly.

Frame: Houston metro (Harris + Fort Bend counties). Sample: 150 anonymized business cohorts, 50 each across 3 sectors — dental, legal, and home services (HVAC, plumbing, roofing). Method: 750 buyer-intent prompts dispatched across 5 frontier assistants — ChatGPT, Claude, Gemini, Grok, and Perplexity — producing 3,750 machine responses. Window: May 1–28, 2026.

Every response was analyzed for three things: was a business named, was the mention carried by a source the user could follow, and were the stated facts true. Those roll into a single 0–100 composite:

Sub-indexWeightWhat it measures
Mention Rate40%Share of relevant buyer-intent prompts in which the business is named at all.
Machine Trust25%Entity completeness a model can verify: schema.org Organization markup, consistent name/address/phone across the web, and a connected sameAs graph.
Citation Capture20%Share of mentions that arrive with a source link the user (and the model) can follow back to the business.
Hallucination Resistance15%One minus the rate at which models fabricate verifiable facts — hours, location, services, ownership — about the business.

The Houston metro median came back at 31 out of 100.

48% of the frame scored below 20 — named in fewer than one in five relevant answers. Effectively invisible. The 90th-percentile floor was 72, which tells you the distribution is not a gentle slope but a cliff with a small plateau on top.

And only 17% of the frame was named by 4 or more of the 5 models on a category prompt. Consensus is rare. Most businesses that exist to one engine do not exist to the others.

The sector split

SectorMedian scoreTop-5 share of all mentionsEffectively invisible
Dental — general, cosmetic, and implant practices38/10061%41%
Legal — personal-injury and family-law firms29/10056%49%
Home Services — hvac, plumbing, and roofing24/10064%55%

Home Services is the worst-served sector we measured and the one with the most concentrated winners: five cohorts capture 64% of all mentions, and more than half the field is invisible.

The finding that reframes everything

We ran correlations between mention rate and every signal a local business could plausibly invest in. The result inverts the standard playbook.

SignalCorrelation with mention rate (r)Type
schema.org Organization completeness0.71machine-readable
Consistent NAP across the web0.55machine-readable
Review volume + recency0.44classic
Google Business Profile completeness0.38classic
Classic domain authority (backlinks)0.21classic

Read the top and bottom rows together.

Backlinks — the central currency of search engine optimization, the thing agencies have sold for twenty years, the metric entire businesses are built on measuring — is the weakest predictor of whether an AI assistant names you. schema.org completeness, which most local businesses have never heard of and most agencies treat as a checkbox, is the strongest, by a factor of roughly 3.4x.

That is association, not causation, and the sample is one metro and 3 sectors. But the ordering is stark enough to demand an explanation, and the obvious one is also the correct one.

Retrieval is not ranking

The reflex is to treat AI visibility as search engine optimization with new inputs. Same game, new referee. Produce more content, add more keywords, build more links, wait for the ranking to improve. Every "Complete Guide to Generative Engine Optimization" published in the last eighteen months rests on that assumption.

Google maintains an index. It crawls a document, stores it, and when a query arrives it retrieves and orders candidates from that stored index. Rank is a position in an ordered list. The list exists. Your job is to move up it, and backlinks are how you do that.

A large language model, asked to recommend an air conditioning contractor in Sugar Land, is not consulting a list of air conditioning contractors in Sugar Land. Depending on the engine and configuration, it is generating from parametric memory, retrieving live documents, and reconciling the two into a single confident-sounding answer.

At no point does it ask which business ranks highest. It asks a prior and much harder question:

Which entities can I resolve here with enough confidence to name one in front of a user?

That is entity resolution — deciding whether scattered records across many sources refer to the same real-world thing. It is one of the oldest hard problems in data engineering, and it is unforgiving in exactly the way ranking is not.

Ranking degrades gracefully. Position 4 still gets clicks. Position 11 gets a few.

Entity resolution does not degrade gracefully. Either the system resolves you into a nameable, corroborated entity, or you are not a candidate at all. There is no position 11. There is inside the answer and outside it, and outside is where 98.8% of local businesses currently live.

This is why schema at 0.71 beats backlinks at 0.21. schema.org Organization markup is a machine-readable assertion of identity: this name, this address, this phone, these services, this sameAs graph. It is exactly the input an entity resolution system needs. Backlinks are a popularity signal for a ranking system that is no longer the one making the decision.

You are not competing for a slot. You are competing to be resolvable.

The threshold

Xi Chu and Yupeng Hou published Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems on June 16, 2026. It has been almost entirely ignored by the commercial AI-visibility field, which is a shame, because it contains the most actionable finding published in this category to date.

The authors ran three experiments across three commercial models — GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash — testing how brand identity, product specification, and marketing language affect what gets recommended.

First: under ambiguity, recommendation is winner-take-all. When competing products carried identical specifications, established brands received 100% of recommendations — an Incumbent Advantage Index of 10.0. A conditional monopoly. Not a strong preference. Everything.

That explains the shape of our Houston data far better than any content-quality hypothesis. When a model cannot distinguish between candidates on evidence, it does not distribute recommendations proportionally. It falls back on the most confidently resolvable entity and gives that one the entire answer. Five cohorts taking 64% of home-services mentions is that mechanism, running in a real market.

Second — and this is the part that matters — the monopoly is fragile. That 100% advantage collapsed once a competitor established a rating differential of less than +0.1 stars.

Read that again. The gap between total invisibility and genuine competitiveness was smaller than a rounding error on a review score. The incumbent's dominance was never strength. It was the model's response to an absence of distinguishing evidence. Supply almost any real signal and the monopoly breaks.

The authors also quantified framing: authority-style claims carried a Bias Surplus Value of +0.17 rating points — more than the entire threshold you need to cross.

Third: the window is arithmetic, and it is closing. In multi-brand competition where every brand optimized, individual payoff collapsed from +0.802 to +0.007. Brands that did not participate received zero.

Early movers capture enormous returns. Those returns decay toward nothing as participation saturates. And abstention does not preserve the status quo — it terminates at zero.

The engines do not agree

ModelCitation rateHallucination rate
Perplexity94%2.4%
Gemini68%3%
Claude41%1.9%
Grok33%5.8%
ChatGPT27%4.1%

Citation rate spans 94% to 27% — a 67-point spread across engines answering the same prompts.

The most-used assistant cites the least. ChatGPT carries a source link on barely a quarter of its mentions, which makes its recommendations the hardest for a buyer to verify and the hardest for a business to trace. If you are wondering why your analytics show no AI referral traffic while customers tell you they "found you through ChatGPT," this table is the answer.

Grok fabricates verifiable facts about real businesses 5.8% of the time — hours, certifications, ownership. Claude is the most conservative at 1.9%; it would rather decline to name a provider than guess one.

Hallucination is not an abstract safety concern here. It is a business telling customers your hours are wrong, and neither you nor the customer knowing where the claim came from.

Why last month's win does not carry

Digital Authority Partners tracked citation behavior across a six-week window and found 1,127 unique URLs cited over the full period. Only 119 appeared in all three measurement waves — 10.6% persistence across 28 days. Profound's analysis of roughly 900 newly published marketing pages found a median of 6.81 days to first citation, with 90% of cited pages cited within 37 days.

Fast in, fast out. AI visibility is not a stock. It is a flow. A ranking you earn in traditional search is an asset that depreciates slowly; a citation you earn in an AI answer is closer to a subscription that lapses.

The Five-Engine Audit

Everything above is diagnosis. Here is the instrument — a field-scale version of the Index protocol, same four sub-indices, same weights. No software, no vendor, no budget. About thirty minutes.

The field's central weakness, identified directly in Olivier Martinez's critical survey of GEO practice, is that it prioritizes algorithmic gaming over verifiable method. Most published guidance is unfalsifiable by construction. This protocol is falsifiable. If it is wrong, you will be able to tell.

Setup

Open five sessions: ChatGPT, Claude, Gemini, Grok, and Perplexity. Use logged-out or temporary sessions — personalization and chat history will contaminate the result. You are measuring the default state a stranger encounters.

The four prompt families

Run each once per engine. Twenty prompts total.

  1. Unaided recall — "Who are the best [category] in [city]?"
  2. Constrained recall — "I need a [category] in [city] that handles [specific service] and can come out this week. Who should I call?"
  3. Direct entity check — "What can you tell me about [your exact business name] in [city]?"
  4. Head-to-head — "Compare [your business] and [top competitor] in [city]. Which would you recommend and why?"

Save every response verbatim.

Scoring

Mention Rate — 40 points. Count how many of the 20 responses named you. Divide by 20, multiply by 40.

Machine Trust — 25 points. The only component measured off-platform, and per our data the strongest lever you have. Five checks, 5 points each:

  • schema.org Organization or LocalBusiness markup present and valid
  • Name, address, phone byte-identical across your site, Google Business Profile, and your top three directory listings
  • A sameAs array connecting your profiles
  • Services listed as discrete machine-readable items, not prose
  • Hours, service area, and credentials in markup — not only in images

Citation Capture — 20 points. Of the responses that named you, how many carried a source link back to a page you control or a page about you? Divide by mentions, multiply by 20. Never named? Score 0 and fix Mention Rate first.

Hallucination Resistance — 15 points. Of the responses that described you, how many contained a factual error? Divide errors by descriptions, subtract from 1, multiply by 15.

Sum all four. That is your score on the same scale as the Index.

Interpretation

  • 0–20 — Unresolved. The engines cannot confirm you exist as a distinct entity. More content will not help. 48% of the Houston frame scored here.
  • 21–45 — Partially resolved. You surface on direct-name queries, rarely on unaided recall. The metro median of 31 sits here, as do all 3 sector medians.
  • 46–70 — Resolved, not preferred. The engines know you but lack evidence to prefer you. Per Chu and Hou, this is the band a differential under 0.1 stars decides. Highest leverage on the scale.
  • 71–100 — Preferred entity. You are a default answer. The Houston 90th-percentile floor was 72. Now the job is maintenance.

What this protocol does not do

Twenty responses is a directional read, not a statistical one — the published Index uses 3,750 for that reason. Engine outputs are non-deterministic; treat a five-point move as noise and a twenty-point move as signal. Results are geographically and sector specific. The Machine Trust component is a hand-scored proxy. And this measures visibility, not revenue — Seer Interactive is the better source there.

The Entity Resolution Stack

The layers below are strictly ordered. Each depends on the one beneath it, and work performed out of order produces no measurable movement.

Layer 1 — Identity. Can a machine determine, without ambiguity, that you are one specific real entity? Every authoritative surface must agree on name, address, phone, hours, and service area — byte for byte. A suite number present in one record and absent in another is, to a resolution system, evidence of two possible entities. This is where schema completeness (r = 0.71) and NAP consistency (r = 0.55) live. Highest-yield work available, and almost nobody sells it because it is invisible.

Layer 2 — Corroboration. Do independent sources confirm you exist? Chen, Wang, Chen and Koudas demonstrated AI search's systematic bias toward earned media over brand-owned content; Meltwater's 5.35M-citation analysis puts earned media at 39.5% of all citations. Your website is an assertion. Independent sources are evidence.

Layer 3 — Evidence. Is there a machine-legible reason to prefer you? This is the +0.1-star layer. Credentials, licenses, response times, warranty terms, years in operation. Not adjectives — values. "Trusted" is unresolvable. "Licensed Texas Master Plumber, M-41892, 24-hour emergency dispatch, 11 years" is a set of comparable facts.

Layer 4 — Retrievability. Can the system access the evidence at the moment it answers? Structured data that validates. Crawlable, non-JavaScript-dependent content. Given ChatGPT cites only 27% of the time, this also determines whether you can ever trace a recommendation. Most agencies start here because it looks like technical SEO — executed on a business that has not resolved Layer 1, it produces nothing.

What this is actually worth

Consumer adoption of AI for local discovery went from 6% to 45% in a year. Visibility sits near 1.2% for businesses considerably better resourced than the median. 48% of the businesses we measured are effectively invisible. The threshold separating invisibility from competitiveness has been measured at under 0.1 rating points. The returns to crossing it are near maximum now and decaying.

The uncomfortable part is that none of this rewards spending. It rewards coherence. The businesses winning AI recommendations are not the ones publishing the most content or buying the most links — links are the weakest predictor we measured. They are the ones a machine can resolve, corroborate, and justify preferring, usually because someone did unglamorous work on identity consistency that no customer will ever see.

That is a strange kind of good news. Content and link budgets compound advantages for incumbents. Entity coherence does not. It is cheap, it is fast, and per Chu and Hou it is decided by margins thin enough that a single-location contractor can cross them in a quarter.

Run the audit. Twenty prompts, thirty minutes. Whatever number comes back, it is the first honest read most businesses will ever have gotten on whether they exist inside the answer.

Right now, for about 98.8% of them, they do not.


Methodology and provenance: figures attributed to the Houston AI Visibility Index are Deep AI Solutions' own measurement, produced by our cross-model AI Visibility Scanner. Cohorts are anonymized by sector; no score is attributed to a named third-party business. Figures are the Index's measured results within its stated frame and window — Houston metro (Harris + Fort Bend counties), 3 sectors, May 1–28, 2026 — not universal facts. Full study: [The 2026 Houston AI Visibility Index](/research/houston-ai-visibility-index-2026).

The Five-Engine Audit protocol is published without restriction. Run it, adapt it, teach it, or publish results against it — including results that contradict ours. A measurement standard nobody can falsify is not a standard.

Shayne Beavan

Shayne Beavan

Founder, Deep AI Solutions · Inventor of record, 6 USPTO filings

Shayne Beavan is the founder of Deep AI Solutions and the inventor of record on its six USPTO filings covering the audit engine, territory lock, drift-correction loop, semantic demand graph, and citation influence engine. He builds and operates the platform from Houston.

Cite this report

Deep AI Solutions. "The 1.2% Problem". By Shayne Beavan. Published July 28, 2026. https://deepaisolutions.com/research/the-1-2-percent-problem