Make better decisions about your AI visibility

Find research to help you decide what to measure, what to change on your website and which AI-visibility claims to question.

How we choose and explain research

How a distill is made

  1. We find papers that address AI visibility, ranking, retrieval, and citation behavior.

  2. We examine the methods, results, and limitations, then isolate findings with practical implications.

  3. We publish the distill in plain English, with the supporting evidence and a direct link to the paper.

All research, newest first 45 articles
01
Sep 14, 2026 7 min read

Body-copy rewrites can cut AI answer visibility

In a realistic AI-search benchmark, the common instinct to rewrite page text for LLMs often hurt the odds of being retrieved, reranked, and cited.

arXiv:2602.12187
02
Sep 14, 2026 3 min read

Citation manipulation lost ground in one AI-search lab test

Before buying a citation-boosting rewrite, ask what the claimed improvement was tested against.

arXiv:2609.02964
03
Sep 9, 2026 10 min read

Source-attributed cues can push multiple-choice models off the right answer

How “source said so” cues destabilize correct answers in multiple-choice QA

arXiv:2609.08934
04
Sep 1, 2026 2 min read

Your AI visibility test may be using the wrong language

For accounting software in Estonia, a small ChatGPT study returned local suppliers in Estonian and none in English.

arXiv:2608.30052
05
Sep 1, 2026 9 min read

Rank-only incentives can make content look better to a ranker while drifting away from human quality

When “rank-only” incentives silently degrade what humans call good

arXiv:2608.30466
06
Aug 25, 2026 9 min read

RAG systems can collapse when they start citing their own output

Self-authored retrieval can “lock” RAG into collapsed answers—at scale.

arXiv:2608.22118
07
Aug 20, 2026 2 min read

An AI answer can get attention without sending you a visitor

A seven-day Google experiment gives marketers a reason to separate answer visibility from website traffic.

arXiv:2608.18352
08
Aug 18, 2026 9 min read

Generative search can erode the web it depends on

How extraction siphons the web until it can’t renew itself

arXiv:2608.15896
09
Aug 10, 2026 9 min read

AI assistants miss 85.6% of venues in a complete Bali market census

Most venues are missing from AI answers—what a true census audit reveals

arXiv:2608.07069
10
Jul 28, 2026 9 min read

Citation type predicts who grounded LLMs name — and roster-based visibility misses almost all of it

Where grounded LLMs actually start naming people—and why most “visibility” measurements miss

arXiv:2607.23893
11
Jul 21, 2026 10 min read

DRNoise shows how one plausible false document can knock deep research agents off course

One misleading “direct claim” can derail an agent that’s otherwise correct

arXiv:2607.17291
12
Jul 20, 2026 9 min read

Chinese generative search cites only a small slice of available brand sources, and external quality scores do not predict what surfaces

Why citation coverage is sparse—and what actually predicts which sources get surfaced

arXiv:2607.15771
13
Jul 16, 2026 8 min read

GEO can change citations inside a fixed context, but it doesn’t show durable organic visibility

Why GEO gains don’t translate into durable visibility (and which levers actually hold).

arXiv:2607.14035
14
Jul 16, 2026 9 min read

LLM brand answers are mostly unstable because language changes the signal

Variance components explain why brand answers won’t stabilize—and what to sample instead

arXiv:2607.13304
15
Jul 7, 2026 9 min read

Open-web search answers more questions, but it makes source trust much harder to control

When open-web search boosts coverage, it quietly degrades source trust

arXiv:2607.05217
16
Jun 25, 2026 9 min read

LLMs mostly source “brand reputation” from other people’s pages

What LLMs treat as “brand reputation” is mostly other people’s web pages, not the brands themselves

arXiv:2606.25787
17
Jun 23, 2026 10 min read

AI brand “ownership” is moderately concentrated, but the winner changes by model

Who gets the “top pick” in AI recommendations—and how consistent is it across models?

arXiv:2606.23057
18
Jun 23, 2026 10 min read

English-only AI reputation monitoring misses local champions in multilingual markets

English-language prompts create a measurable “local-visibility” blind spot across languages

arXiv:2606.23165
19
Jun 23, 2026 9 min read

AI visibility breaks down by entity, not just by mention count

Why “mention” counts fail: fabricated citations scale differently by entity and query context

arXiv:2606.21595
20
Jun 19, 2026 9 min read

AI search visibility starts with brand stature, not prompt tweaks

Why the first-run visibility gap between big brands and everyone else is so persistent

arXiv:2606.20065
21
Jun 17, 2026 9 min read

Incumbent brands get a built-in advantage in LLM recommendations — but only until a competitor clears a narrow threshold

How LLM recommenders lock in incumbent brands—and the small tweaks that break it

arXiv:2606.17443
22
Jun 16, 2026 10 min read

LLM search agents can be pushed to endorse manipulated web content

When “search-and-answer” becomes endorsement for sale

arXiv:2606.16821
23
Jun 12, 2026 9 min read

One Polluted Page Is Enough to Hijack LLM Recommendations

Why one poisoned search result is enough to hijack LLM recommendations

arXiv:2606.13610
24
Jun 10, 2026 10 min read

AI brand recommendations move people onto the open web through search, not just mention counts

When an assistant says the brand name: what actually moves browsing

arXiv:2606.10907
25
Jun 9, 2026 9 min read

Safety-trained RAG models can turn a prompt injection into brand suppression

Why safety alignment can turn retrieval-time injections into brand-level anti-promoters

arXiv:2606.09204
26
Jun 8, 2026 8 min read

FullCite gets better quote grounding by separating the document from the evidence span

FullCite turns inline citation into a document-plus-span problem, sharply improving quote-level grounding on ASQA.

arXiv:2606.07130
27
Jun 4, 2026 10 min read

ChatGPT referral spikes can overstate AEO unless you control for platform growth

Why “2x on ChatGPT” stories can be misleading without a tailwind control

arXiv:2606.04362
28
Jun 2, 2026 7 min read

LLMs Score 94% on Cultural Knowledge Tests and 40% When the Answer Choices Are Removed

When Cultural Knowledge Doesn't Transfer to Cultural Reasoning

arXiv:2606.01879
29
Jun 2, 2026 7 min read

English Prompts Suppress Bengali Cultural Knowledge Even When Local Evidence Is Provided

How Prompt Language Rewrites Cultural Knowledge Before the Model Even Answers

arXiv:2605.30481
30
Jun 2, 2026 7 min read

LLM Fact-Checkers Score Well But Retrieve the Wrong Sources

Where LLM Fact-Checkers Go Wrong on Sources

arXiv:2605.30241
31
Jun 2, 2026 7 min read

An AI agent writes better-rated health notes by learning from past corrections

A GPT-4.1-based judge favored generated health notes; the score does not measure real-world correction success.

arXiv:2606.02215
32
Jun 1, 2026 7 min read

AI Overviews Sent Users to Reddit. AI Mode Is Taking Them Back.

Google AI Overviews drove a 12% rise in Reddit engagement, but AI Mode reversed those gains for experiential communities by substituting conversation for human discussion.

arXiv:2605.16428
33
Jun 1, 2026 7 min read

Ecosystem GEO Beats Page-Level Optimization by Up to 31 Points for Agent Search

Coordinating a multi-page evidence ecosystem raises LLM search agent recommendation rates by up to 31 percentage points over the best single-page GEO baseline.

arXiv:2605.12887
34
Jun 1, 2026 7 min read

When AI Cites AI: The Synthetic Source Problem in Generative Search

An audit of ChatGPT, Copilot, Gemini, and Perplexity finds ~16% of cited sources are AI-generated — with Copilot citing synthetic content in nearly 3 of every 10 citations.

arXiv:2605.23684
35
Jun 1, 2026 6 min read

Frontier LLMs Hallucinate Up to 38% of Scientific Citations

Six frontier LLMs hallucinate 12–38% of scientific citations; a new agentic retrieval system hits zero hallucination at 30% better F1 and $0.05 per query.

arXiv:2605.14306
36
Jun 1, 2026 6 min read

Rewording a Buying Question Changes the Brands AI Recommends More Than Switching Models Does

Cosmetic prompt rewording drops AI brand-recommendation overlap by 21–32 percentage points — more divergence than switching providers entirely, across 12,000 runs.

arXiv:2605.27440
37
Jun 1, 2026 6 min read

Query-Specific Expiry: Why 'Recent' Isn't the Same as 'Fresh'

Baidu's Aurora-Expiry uses RAG-augmented LLMs to infer query-specific expiration thresholds, cutting median document age 12.81% for time-sensitive queries in a 14-day live A/B test.

arXiv:2605.13052
38
Jun 1, 2026 8 min read

RAG Doesn't Flatten the Brand Hierarchy — It Just Moves Where You Lose

A 37,000-run audit of 533 brands finds RAG preserves the brand hierarchy: L4–L5 specialists face 48–52% invisibility while L1 leaders surface universally but convert at only 25–41%.

arXiv:2605.27439
39
May 31, 2026 7 min read

Semantic Metadata Makes Agents More Reliable, Not Smarter

Schema.org markup gives retrieval agents 65.7% higher FAIR-compliant precision — but cuts query coverage by 29% where publishers haven't adopted it.

arXiv:2605.28787
40
May 26, 2026 6 min read

One false search result sharply reduced AI accuracy in a synthetic web test

A Microsoft study shows a single false top search result drops GPT-5 accuracy from 65% to 18% — while humans solve the same queries at 93% — exposing a critical gap in agentic RAG deployments.

arXiv:2603.00801
41
May 22, 2026 8 min read

AI Research Agents Fail in Ways Their Own Tests Can't See

A new SoK survey of 118 works shows that agentic RAG's iterative retrieval and memory systems introduce failure modes that static metrics and current benchmarks cannot detect.

arXiv:2603.07379
42
May 22, 2026 7 min read

Most Pages Get Zero AI Citations. Editing 5% Won 40% More.

A new paper from Virginia Tech maps four failure modes that prevent pages from being cited in AI-generated responses. 43% of relevant pages receive zero citations under baseline conditions.

arXiv:2603.09296
43
May 22, 2026 7 min read

AI Cites Sources It Never Checked — and Half Don't Hold Up

No LLM verifies even half its citations under any tested condition — and adding temporal cutoffs or other deployment constraints collapses verifiability to near zero.

arXiv:2603.07287
44
May 22, 2026 8 min read

Generative Search Citation Share Is a Noisy Estimator, Not a Score

A new statistical framework shows that single-run citation share metrics from Perplexity, SearchGPT, and Gemini carry confidence intervals wide enough to make most apparent SEO gains statistically indistinguishable from noise.

arXiv:2603.08924
45
May 22, 2026 2 min read

Before paying for more schema, check what the AI can read

One retrieval experiment favored richer, clearer entity pages. It does not show that schema is useless in public search.

arXiv:2603.10700