Open AI discovery intelligence

Understand how machines discover the web.

The open intelligence layer for AI discovery: a living, evidence-backed record of how machines discover, retrieve, cite, recommend, and act on information.

6latest material changes
25controlled topics
25registered sources
5evidence claims
Latest intelligenceView full timeline →

What changed, and why it matters.

Every update is placed inside a persistent historical, semantic and evidentiary record instead of disappearing into a news archive.

research

Perplexity Q2D-Web maps the retrieval stage behind agentic search

Nearly 70,000 machine-reformulated queries show why the query a user types is not necessarily the query your content competes for.

PerplexityRead evidence record →
major

Cloudflare’s new Search / Agent / Training crawler defaults take effect

Cloudflare’s announced defaults take effect for new domains: Training and Agent traffic is blocked by default on ad-supported pages while Search remains allowed; multi-purpose crawlers are governed by their most restrictive applicable purpose.

Event recordOpen →
major

Perplexity introduces Q2D-Web for large-scale retrieval evaluation in agentic RAG

Perplexity released Q2D-Web, a production-shaped retrieval benchmark built around 190 million web documents and 69,721 agent-reformulated queries across 10 languages.

Event recordOpen →
notable

Cloudflare expands BotBase for bot and agent transparency

Cloudflare expanded BotBase so bot and agent operators can identify themselves and document what their automated systems do.

Event recordOpen →
major

Cloudflare separates AI traffic into Search, Agent, and Training

Cloudflare introduced a crawler-purpose taxonomy separating Search, Agent and Training traffic and announced new default controls.

Event recordOpen →
major

Google publishes official guidance for generative AI Search visibility

Google Search Central published a guide explaining generative AI Search, including RAG, query fan-out, and its position that core SEO practices remain relevant.

Event recordOpen →
Narrative architectureExplore history →

AI visibility did not begin with GEO.

Visibility OS follows the longer story: how documents became machine-readable, how search learned entities, how answers replaced result lists, and how agents now fan out queries across the web.

1

Semantic web

Structured data made meaning more explicit to machines.

2

Entity understanding

Search moved from strings toward entities and relationships.

3

Answer extraction

Featured answers made being selected different from simply ranking.

4

Retrieval-grounded generation

RAG connected generative models to external information.

5

Generative search

AI systems began synthesizing answers across multiple sources.

6

Query fan-out & agents

A single prompt can trigger many machine-generated searches and actions.

Ask the corpusSearch everything →

Explore by question, not just category.

The interface should match how people actually investigate AI discovery: by asking what changed, what is proven, and what remains uncertain.

Semantic architectureBrowse all topics →

Living reference pages, not disposable posts.

AEO, GEO and LLMO sit inside a broader graph of retrieval, source selection, citations, discoverability, measurement and agentic behavior.

Evidence disciplineBrowse claims →

Facts, interpretation and uncertainty stay separate.

Each claim is attached to source quality and evidence status. Visibility OS is designed to preserve the evidence trail, not flatten every industry claim into advice.

EstablishedAuthoritative or directly verifiable
StrongSubstantial support with limited uncertainty
EmergingSupported, but still developing
Mixed / OpenConflicting evidence or unanswered questions
Open questionsResearch library →

What the industry still does not know.

Uncertainty is part of the dataset. We track unresolved terminology, causal claims and measurement problems instead of pretending consensus exists.

Research gap 01

What is the earliest verifiable published use of the term “Answer Engine Optimization” (AEO)?

The practice predates the terminology, and current web sources frequently repeat unattributed origin claims.

Research gap 02

What is the earliest verifiable use of “Large Language Model Optimization” / LLMO for web visibility?

The acronym is used inconsistently and overlaps with unrelated ML optimization meanings.

Research gap 03

Does llms.txt materially affect inclusion, citation frequency, or retrieval in major AI search systems?

Adoption does not establish causal visibility impact.

About the dataset

Built to be read by people and machines.

Semantic HTML, explicit dates, controlled vocabulary, source attribution and stable canonical pages make each record understandable even when the visual layer is stripped away.

Discovery can be automated, but publication is reviewed. Important records retain primary sources, evidence notes, relationships and revision history.

Original sources and research records are also archived separately so the intelligence corpus is not dependent on one database or hosting provider.