Sunday, October 4, 2026 Five things that moved Read time: 8 min

Today’s theme · Which sources AI actually names

ChatGPT and Gemini cited no source in common on 77% of the same questions

Four audits published in the last three weeks measured which sources AI assistants name. The assistants do not agree with each other, with the public record, or with themselves.

Four research teams published audits in the last three weeks that ask a narrow, commercial question. When someone asks an AI assistant for a recommendation, which sources does it actually name?

Two of the audits arrived at the same place from different directions. The two most-used assistants rarely cite the same sources as each other, and a single assistant rarely cites the same sources twice in a row.

A third tested the tactic most often sold as the fix for this, and found it moves credit between sources without reliably deciding which sources get in at all. A fourth found that the wording of the question changes what comes back.

None of this says AI search can be ignored. It says the thing being measured is noisier than the people selling the measurement usually admit, which is worth knowing before you buy any of it.


01 Research & evidence

ChatGPT and Gemini shared 5.4% of their sources

Being named by one assistant tells you almost nothing about the others, so there is no single place to be found.

On 76.7% of the identical questions put to both assistants, the two had not one source domain in common.

A team led by Lucas Uberti-Bona Marin curated ConsumerQ, a set of 2,528 real commercial-advice queries, then evaluated 1,536 responses to product questions from ChatGPT, Google Gemini and Google AI Overviews, covering both the consumer chatbots and the developer APIs.

The headline number is the disagreement. For the same question, the ChatGPT and Gemini interfaces shared only 5.4% of their source domains on average, and on 76.7% of comparisons they shared no domain at all.

The developer APIs did not stand in for the chatbots either. Mean domain overlap was 12.0% for ChatGPT and 14.8% for Gemini between an API and its own consumer interface, and the two expose different layers of source information. Anyone measuring AI visibility through an API is measuring something a customer never sees.

One more finding belongs on the record. ChatGPT stated a first-person product preference in 79% of its product-recommending responses, against 7% for Gemini and 2% for AI Overviews, and the products it named often changed between repeated identical requests.

What to do about it

Ask your own buying question in ChatGPT and in Gemini, in the ordinary consumer apps rather than through any tool, and write down which sources each one shows. Then ask both again tomorrow.

Two lists that barely overlap, and that move between days, is the normal result. If your business sells through comparison, an insurance brokerage or an admissions consultancy as much as a software company, the practical read is that there is no one assistant to win.

Source arXiv, Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig, “Auditing AI-Generated Product Recommendations,” September 16, 2026 · arxiv.org

02 Research & evidence

AI fabricated most provider referrals with search off

An assistant answering about your field without searching the web may name a business that does not exist in your city.

Switching retrieval on took the same model from 11% real referrals to 71%, so the setting mattered more than the brand of AI.

Hazem Ibrahim and Yasir Zaki audited AI provider recommendations across the 100 largest United States metropolitan areas in four service domains where an official registry exists, matching every recommendation against Medicare clinician and facility records and SEC adviser disclosures.

With web search off, the recommendations largely were not real. Only 4% of the open-weight model’s recommended doctors and 11% of the proprietary model’s matched a clinician in the city asked about, and the open-weight matches were name coincidences: those clinicians were no likelier to be primary-care doctors than names drawn at random from the registry.

Turning retrieval on changed the answer. With search, 64% to 71% of recommendations in the same domains matched a real provider. Search also removed the metro-size penalty, so smaller markets stopped being systematically worse served.

The finding for wealth advisors and RIAs is the one to carry into a meeting. Without search, recommended advisory firms carried SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size, and with search they fell significantly below it.

The authors are blunt about why this is hard to catch. An answer produced without retrieval carries no sign that its recommendations were never verified.

What to do about it

Ask an assistant the question a prospective client would ask in your field and your city, once with web search clearly on and once in a plain chat window. Compare the names against your own licensing board or registry.

For physicians, surgical practices and advisory firms this is a referral, whatever the interface calls it. A profile a retrieval system can find and verify is what moves an answer out of the fabricated column.

Source arXiv, Ibrahim and Zaki, “Understanding AI Provider Recommendations in Local Service Markets,” September 16, 2026 · arxiv.org

03 Research & evidence

Formatting a page moved credit, not whether it is cited

Structuring a page changes how much credit it gets once an assistant picks it, and not whether the assistant picks it.

The effect on being cited at all was 4.5 percentage points and did not clear the study’s own pre-specified bar.

Sriram Selvam and Anneswa Ghosh built CITECHOICE to test the assumption most content advice in this field rests on. When several retrieved sources support the same claim, how does an answer engine choose which to cite? They took 129 real multi-turn search transcripts, selected 113 document pairs that independently supported the same fact, and had 103 of them confirmed by blinded human review.

Then they replayed each transcript with one source rendered as structured content and as plain prose, holding everything else fixed. Structured rendering raised that source’s citation count by 0.50 citations per answer, with a 95% confidence interval of 0.20 to 0.84.

The effect everyone actually wants was the one that did not land. Whether the source was cited at all moved 4.5 percentage points, on a confidence interval running from minus 1.4 to plus 10.4, which the authors record as inconclusive. Formatting concentrated credit among sources already in contention rather than letting new ones in.

Position still dwarfed presentation. The citation-rate gap between rank 1 and rank 5 was 42.3 percentage points, against 7.9 points from the controlled reordering the study ran itself.

Structured markup is one of the things every firm in this field sells, Forever Cited included. A causal audit finding that the effect sits on how much credit a source gets, rather than on whether it gets in, is worth stating plainly by the people with an interest in the other answer.

The study also measured its own noise. Re-running the same prompts changed 15% of the cited-or-not decisions, and decoding randomness accounted for an estimated 45% of the variance in single-generation effects.

What to do about it

Keep structuring your pages. The evidence says it helps once a source is in contention, which is a real effect and a modest one.

What the evidence does not support is a claim that formatting alone decides whether an assistant uses you. If anyone quotes you a citation rate that moved on formatting, ask how many times they ran it, because 15% of those answers changed on a re-run.

Source arXiv, Selvam and Ghosh, “CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search,” September 14, 2026 · arxiv.org

04 Research & evidence

Retrievers returned worse health sources for one dialect

The words a customer uses to ask change which sources come back, so your visibility is not one number.

Every system tested did worse on consumer-health questions written in African American Language than on the same questions in mainstream English.

Andrew Tang, Nicholas Deas, Kathleen McKeown and Vishal Misra tested whether the identity signals people carry in their own phrasing change what a search system retrieves. They built paired evaluations in political news and consumer-health questions, varying only that signal, and ran them across five dense retrievers and a sparse baseline.

Both results held everywhere. Every retriever returned articles aligned with the political lean of the question itself, and every retriever performed worse on questions written in African American Language than on the same questions in White Mainstream English.

The authors checked the obvious explanation and ruled it out. Removing an aggregate measure of word-choice differences left the gaps largely intact, and simple probes recovered both the lean and the dialect from the system’s internal representation of the query, beyond the individual words used.

For a business this is not a story about politics. It is evidence that there is no single retrieval result for your category, because the customer’s own phrasing is part of the query.

What to do about it

Write for the range of ways your customers actually ask, not the phrasing you would use in a brochure. A clinic, an insurance brokerage and an architecture practice all have customers who describe the same need in different registers.

If you measure AI visibility, measure it on more than one phrasing of the same question. A single wording is a single sample, and this study is a direct demonstration that the sample moves.

Source arXiv, Tang, Deas, McKeown and Misra, “Retrieval Sensitivity to Identity Signals in Queries,” September 29, 2026 · arxiv.org

05 Platform change

Google Local Service Ads are hiding the phone number

If you pay per lead on Google, the number a customer needs may now sit behind an extra click.

Google has confirmed nothing, so treat this as a sighting to check on your own ads rather than an announced change.

Barry Schwartz reported on October 2 that Google Local Service Ads are showing a “Get phone number” button where the number used to sit, with the number appearing as soon as a cursor moves near the ad.

The reporting is honest about its own limits. Schwartz writes that he is not sure whether this is new, and Google has not confirmed it, announced it, or described it as a test. One observer, one day, no statement.

It is still worth two minutes of your time. Local Service Ads are the pay-per-lead placements that run above ordinary search results for licensed local services, which means the phone call is the product you are buying.

The desktop behavior and the phone behavior are not the same thing. A hover reveals the number on a desktop, and there is no hover on a phone, so any effect on call volume would land hardest on mobile searches.

What to do about it

Open a private window, search the term you buy Local Service Ads on, and look at your own ad on a phone as well as a desktop. If the number sits behind a button, your cost per call may move before anyone tells you why.

Do not change your bidding on one sighting. Note what your ad looks like today, so that if your call volume moves next week you can tell a layout change from a demand change.

Source Search Engine Roundtable, Barry Schwartz, “Google Local Service Ads Hiding Phone Number By Default,” October 2, 2026 · seroundtable.com

If you do one thing this week

Ask the question a prospective customer would ask about your field, in ChatGPT and in Gemini, in the ordinary consumer apps. Write down the sources each one names, then do it again in two days.

Four separate teams measured versions of that comparison this month and all four found instability: 5.4% average source overlap between the two assistants, 4% to 11% of provider referrals matching a real local clinician without retrieval, and a formatting effect that moves credit without reliably deciding who gets in.

The useful conclusion is not that any of this is hopeless. It is that a single check, on a single day, in a single wording, is not a measurement, and anyone quoting you one as though it were has not run it twice.

Read the full archive →

Forever Cited · AI Visibility, Engineered by Industry · One brand per market.
Sources are linked in full above. We link to primary documents and original research wherever they exist.

The daily brief

Get this briefing every morning — free.

Five sourced items on AI search and visibility, written for business owners. No pitch, no spam. Confirm by email (double opt-in); unsubscribe anytime.

← All AI Industry Updates  ·  forevercited.ai