# Profound’s AEO research: a reading list with the disagreements left in

Canonical: https://mudpie.ai/blog/profound-aeo-research-reading-list/
Breadcrumb: [Home](https://mudpie.ai/) / [Field notes](https://mudpie.ai/blog/) / [Profound’s AEO research: a reading list with the disagreements left in](https://mudpie.ai/blog/profound-aeo-research-reading-list/)
Author: Ali Abouelatta (https://mudpie.ai/authors/ali-abouelatta/)
Published: 2026-09-19
Updated: 2026-09-19
Research type: research-reading-list
Method: Read the full Profound research pages, preserve each study’s own denominator and date range, and separate vendor observations from independent validation and causal evidence.

Profound’s research archive is useful when you read it like a set of studies. It becomes misleading when you read it like one continuous proof that every visibility problem has one answer.

The archive mixes prompt-intent analysis, citation-source analysis, volatility, HTML structure, and product methodology. Those are different jobs. The numbers do not share one denominator, one date range, or one definition of “visibility.”

## Start here: the reading list

| Read | Reported result | What the method gives you | What it cannot settle |
|---|---|---|---|
| [AI Search Volatility](https://www.tryprofound.com/blog/ai-search-volatility) | About 80,000 prompts per platform, comparing June 11–13 with July 11–13, 2025. Domain drift: 59.3% Google AI Overviews, 54.1% ChatGPT, 53.4% Copilot, 40.5% Perplexity. | A concrete warning against one-run scores. | The prompt list, sampling frame, and model versions are not public. Drift means domains changed, not that every brand disappeared. |
| [Citation Overlap Strategy](https://www.tryprofound.com/blog/citation-overlap-strategy) | 100,000 prompts across ChatGPT and Perplexity: 37.4% ChatGPT-only citations, 51.6% Perplexity-only, 11.0% overlap. | A strong argument for engine-specific measurement. | The page does not publish the prompt construction, geography, or exact run design. |
| [AI Search Shift](https://www.tryprofound.com/blog/ai-search-shift) | From a 240M-citation dataset, ChatGPT–Google alignment rose 12% in April 2025 to 33% in July; ChatGPT–Bing fell 26% to 8%. In a 1,000-prompt overlap subset, Google position one drew 10% of ChatGPT citations versus 27.5% of human clicks. | A useful bridge between SEO rank and AI source selection. | The position comparison applies only when ChatGPT and Google sources align. Most ChatGPT sources remained independent of Google’s top results. |
| [AI Platform Citation Patterns](https://www.tryprofound.com/blog/ai-platform-citation-patterns) | 680M citations from August 2024–June 2025. ChatGPT’s largest overall source was Wikipedia at 7.8%; Google AI Overviews’ Reddit share was 2.2%; Perplexity’s Reddit share was 6.6%. | A cross-engine source-stack map. | Overall citation share and “share inside the top ten” are different denominators. The public article does not disclose prompt count or raw sampling. |
| [AI Search Intent Study](https://www.tryprofound.com/blog/chatgpt-intent-landmark-study) | A sample described as tens of millions of ChatGPT interactions: generative 37.5%, informational 32.7%, no intent 12.1%, commercial 9.5%, transactional 6.1%, navigational 2.1%. | A reason to question SEO-derived prompt lists. | Exact n, geography, collection window, and classifier validation are not disclosed. The comparison to traditional-search intent is not apples-to-apples. |
| [2,000-page structural study](https://www.tryprofound.com/articles/how-to-optimize-answer-engines-2025) | Profound says thousands of pages across hundreds of domains favored tables, numbered headings, FAQ/HowTo schema, detailed alt text, and 1,500–2,500-word pages. | A source of page-structure hypotheses. | No absolute lift, public control, query list, or confidence interval. “Significantly” is not a number here. |
| [The Data on Reddit and AI Search](https://www.tryprofound.com/blog/the-data-on-reddit-and-ai-search) | More than 4B AI citations and 300M answer-engine responses; Reddit’s aggregate citation share was 3.11% from August 2024 to late October 2025; the average cited post was about one year old. | A case for treating human conversation as part of the source stack. | Profound worked with Reddit; the public page does not disclose the full sampling frame or classifier validation. |

## Where the studies disagree

### 1. A large dataset does not remove sampling uncertainty

The volatility study is the most useful place to start because it changes how the other numbers should be read. If 40%–60% of cited domains change over a month for repeated prompts, then a visibility percentage is a sample from a moving distribution. A larger sample reduces random noise inside a given collection window. It does not freeze the engine, the web, or the prompt population.

That is consistent with a [2026 preprint on AI-visibility uncertainty](https://arxiv.org/abs/2603.08924). It sampled three platforms across three consumer-product topics, with daily collections for nine days and high-frequency collections every ten minutes. Many apparent domain differences fell inside the reported noise floor. It is a small-topic study, not a universal estimate of engine variance.

### 2. “Prompt intent” is not one clean market census

The 50M+ study is the most shareable Profound number and the least complete method description. “Generative” at 37.5% is a reported label in a vendor taxonomy, not a directly observable category like a Google Search Console query type. The page does not state how the sample was selected, whether repeated users or conversations were deduplicated, or how the labels were validated.

Profound’s [Prompt Research Reports page](https://www.tryprofound.com/features/prompt-volumes/research-reports) describes retrieval, ranking, clustering and coverage-based selection over a claimed 1.9B+ conversations from six months. It is a vendor description without public engine-by-engine counts or an external audit. The [earlier announcement](https://www.tryprofound.com/blog/introducing-prompt-research-reports-in-profound) said 1.5B+ prompts. Those changing totals and units should travel with the date and source, not be treated as one fixed population.

### 3. Cross-engine source shares cannot be merged into one “AI internet”

Profound reports ChatGPT, Google AI Overviews, and Perplexity leaning on different sources. Its overlap study finds only 11% shared domain citations between ChatGPT and Perplexity in a 100,000-prompt sample. The platform-citation report then gives different source shares for each engine.

That makes “AI visibility” a bad singular noun unless the article names the engine. A brand can be visible on Google AI Overviews and absent on ChatGPT. A page can be valuable for Perplexity and irrelevant to Copilot. The practical reading is platform-specific source strategy, not a universal citation ranking.

### 4. Structure is an observed pattern, not a guaranteed lever

Profound’s 2,000-page article and the similar AirOps structure study point toward tables, headings, schema, lists, and clear phrasing. Google’s own guidance says no special schema is required for AI Overviews or AI Mode. The original academic GEO paper found that citations, quotations, and statistics changed answer composition in a fixed five-source GPT-3.5 simulator, with the best internal metric improving by 41%.

Those findings can all be true. They answer different questions:

| Question | Evidence strength |
|---|---|
| Is clean, indexable, people-first content still required? | Strong official guidance from Google. |
| Do structured pages correlate with being cited? | Directional vendor/agency evidence. |
| Can text edits change how much of an already-retrieved source appears in an answer? | Original academic benchmark evidence. |
| Does adding FAQ schema cause more live citations or revenue? | Not established by these studies. |

## The reading order I would use

Read volatility first. Then read overlap and platform citation patterns. Then read the intent study with its sampling gaps in view. Read the structural article last, as a list of hypotheses. The Reddit study is useful for source-type analysis, but keep the Reddit partnership and the changing time window in the paragraph.

For an editorial briefing, cite the volatility study first, then attach the intent study’s sampling note to every prompt-volume number. Read the structural article as a list of hypotheses to test.


## Author disclosure

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.
