Skip to Content
DocsMonitoringAI Sources
Monitored dimension

AI Sources

What it tells you

Being recommended and being read are two different facts, and only one of them is something you can go do something about this week. When an engine answers “what’s the best tool for X”, it pulls a handful of pages first — review roundups, comparison posts, category directories, competitors’ own sites — and writes its answer from those. AI Sources shows you that reading list.

What you seeWhat it means
The coreThe hosts that more than one engine opened. The pages your market’s answers are actually built from.
Your own hostsCounted in the core total, and excluded from the funnel — so the two core numbers differ by your own sites, deliberately.
Where you’re already namedCore hosts whose pages mention you. Won — no work needed.
The work listCore hosts that were read, where no page names you. This is the actionable part.
Competitor-ownedCore hosts a rival owns. Real, and not something outreach fixes.
Your own pages, when openedAn engine opened your site, then named these companies. Being read is not being recommended.

Retrieved, never cited. An engine hands back the pages it opened while answering; it never says which ones it leaned on. No number on this page is a citation count.

The engines are never pooled. Perplexity and Google AI Overviews read different pages, so every page count belongs to one engine and is never added to the other’s. The one shared object is the core — hosts that at least two engines opened.

Counts are counts, never rates — on screen. The app reports 3 of 8 answers rather than 38%, because eight questions is a deliberately small set and a bare percentage computed from it is false precision. The API does ship a share, but never bare: presence always travels with presenceLow and presenceHigh, so the width of the range is part of the reading.

What a check looks like

A check asks each engine your 8 buying questions, reads the companies named in each answer, and collects every page the engine opened along the way. One run reads something like this:

ai sources · one checkillustrative — 8 questions × 2 engines
perplexity 8 asked · 8 answered · 47 pages opened
google_ai_overviews 8 asked · 6 answered · 2 no answer shown · 21 pages opened
├ core hosts (both engines) 14
│ ├ already naming you 5
│ ├ could not be read 2
│ └ missing you 7 · 4 publishers · 3 competitor-owned
└ verdict missing_from_most_core_hosts
your pages opened in 2 answers — both named 5 other companies

Every count names its universe on the same object, so you always read them in pairs: answers naming you of answers received, independent pages naming you over pages read. A shortfall is two facts rather than a ratio — “8 asked, 6 answered” — because the reasons differ and one of them is ours.

The funnel is the finding

The dimension narrows in four steps, and reading it in order is what keeps it honest:

  1. Third-party and competitor hosts both engines read — the core.
  2. Those already naming you — won, and removed from the list.
  3. Those you’re genuinely missing from — with the hosts we couldn’t read at all noted underneath rather than counted in, since a host nobody could read is not one you are absent from.
  4. Those you can approach — step 3 minus the competitor-owned sites, which stay visible below it.

On a strong brand step 3 is small, because the core already names you. That reads as “already on 19 of the 25 hosts both engines read” — not as “nothing found”. An empty work list is a result, not a failure to find one.

One number that catches people: unreadable at host level and unreadable at page level are different universes. A host only counts as unreadable when no page on it could be read, so a check can report zero unreadable hosts while listing twenty-five unreadable pages.

The work list, and what isn’t on it

The work list is every core host whose status is missing — at least one page there was read and none of them names you. Two statuses are deliberately excluded:

StatusWhy it isn’t work
already_namedYou’re on it.
unreadableNo page on it could be read — so it is not a host you are absent from.

Competitor-owned sites are on the work list, not excluded from it. Ownership is a separate field from status, so a rival’s own site that was read and doesn’t name you is missing like any other. It is kept and counted, and shown as a second group beneath the approachable hosts — context, never a target. That split matters more than it sounds: a customer can have fourteen hosts missing them of which thirteen are competitor-owned, and reading that as one row of work would badly understate the position.

Pages we couldn’t read carry a reason rather than a blank — nine of them, covering the ordinary failures (unreachable, timeout, fetch_failed, unresolved, bot_protection, host_blocked_recently, text_too_short) and two that look like successes: not_found_page, where the server answered normally with a page saying the resource doesn’t exist, and consent_wall, where a consent interstitial stood in place of the content. Whichever it is, the page is listed and enters no count.

Each host carries an actionHint with ready-made wording. In the app and over the API that text is rendered as-is rather than paraphrased — it’s written to match the host’s status, and a sentence composed from the status code alone tends to overclaim.

The verdict is a condition, not a grade

Each check stores one of three verdicts, decided when the summary was built and never re-derived from the numbers:

VerdictWhat it means
recommended_nowhereNo answer in the window named you at all.
named_on_most_core_hostsYou’re on at least half the core — so a short work list is the finding, not an empty result.
missing_from_most_core_hostsNamed somewhere, absent from most of the hosts more than one engine read.

It’s a condition code rather than a score, so it’s always read beside the counts it rests on. “Seen in N of M” figures are per engine, over a window of up to five published checks that asked the same engine roster. On a first check they don’t appear at all — rather than print “1 of 1”, which would read as a measurement, the app hides the column and says the figures arrive from your second check.

Three ways an answer can go

A question that didn’t produce an answer naming you has not necessarily produced a fact about you. The dimension keeps three states apart, and collapsing any two of them misreports the measurement:

StateWhat happenedHow it’s reported
AnsweredThe engine answered.An answer that recommended nobody is a real finding, not a gap.
No answer shownThe engine was read and showed nothing — Google’s results page carried no AI Overview.Measured, not a failure. In no denominator, and never “not named”.
Not measuredWe couldn’t read the answer.Our problem. In no count, and never “not mentioned”.

An answer that opened no pages is reported as “no pages reported” rather than “answered from memory” — the engine hasn’t said how it answered, and we don’t guess.

The funnel, the work list and the per-host hints live in the CompetLab app. The REST API and MCP tools return the same stored summary, plus the engines’ raw answers and the pages each opened, so you can rebuild the view in your own tools.

A trend, not a snapshot

One check tells you today’s reading list; the value builds as it repeats. Because the pooled figures are read over a window of published checks, a host that keeps appearing is a stronger target than one that surfaced once — and a host that stops appearing tells you the market’s reading moved. When the engine roster changes, the pooled window restarts rather than mixing two rosters together, so a change in what we asked never shows up as a change in the market.

AI Sources runs on its own cadence, set independently of the other dimensions, with a minimum of three days between checks — the same floor AI Visibility has. Turning it on and choosing how often it checks are both covered in How monitoring works.

Work with it in code

FAQ

What is AI Sources?

AI Sources is one of CompetLab's six continuously monitored dimensions. It shows you which pages AI engines actually open when they answer your buyers' questions, and whether your site is among them. Each check puts your eight buying questions to Perplexity and Google AI Overviews, records the companies each engine named, and collects every page the engine opened to write its answer. Those pages are then grouped by host, so you can see which sites your market's AI answers are built from — the review roundups, comparison posts and directories the engines keep returning to — and which of them name your competitors and not you. The result is a work list of sites where being absent is costing you, rather than a score. Because it's monitored, every check is kept as history you can read as a trend. It is the one monitored dimension that does not raise alerts yet — the work list is re-derived on each check instead.

How is it different from AI Visibility?

They answer two different questions about the same moment. AI Visibility asks whether the engines recommend you — how often you're named, against which competitors. AI Sources asks what those engines read before deciding, and whether your site is in that reading. You can score well on one and badly on the other: a brand can be recommended often while being absent from every page the engines open, or be widely read and rarely named. The two are meant to be used together, which is why they sit side by side in the product — the first tells you where you stand, the second usually tells you why.

Does AI Sources tell me which pages the engines cited?

No, and that distinction matters. An engine hands back the pages it retrieved while answering; it does not disclose which of them it actually leaned on to write the answer. So every number here is a count of pages retrieved, never of pages cited, and we say "retrieved" deliberately throughout. In practice a retrieved page is still a strong signal — an engine opened it while forming an answer about your category — but treating it as a citation would claim more than the data supports. For the same reason, page counts are never added across engines: Perplexity and Google AI Overviews read different pages, so each engine's count belongs to that engine alone.

Which engines does it use, and how many questions does it ask?

Two engines — Perplexity and Google AI Overviews — and eight buying questions per project. Those two are the engines that answer by searching and hand back the pages they opened, which is the whole basis of this dimension; the wider engine roster lives in AI Visibility, where the fact being measured is who gets recommended rather than what was read. Eight questions is a deliberately small, high-intent set, which is why every figure is reported as a count rather than a percentage: a share computed from eight answers reads far more precisely than it deserves. Pooled figures are read over a window of up to five published checks that asked the same engine roster.

What do I actually do with the work list?

The work list is the set of hosts that more than one engine opened, where at least one page was read and none of them names you. Those are the pages your buyers' AI answers are being written from, so getting named on them is the concrete move — a review roundup you're missing from, a comparison post that lists your rivals, a category directory you never submitted to. Two things are deliberately kept off the list: hosts that already name you, and hosts where no page could be read, since a host nobody could read is not one you are absent from. Competitor-owned sites stay ON the list — they are grouped separately beneath the approachable hosts and marked as context rather than targets, because a rival's own page naming your rivals and not you is a real reading even though outreach will never win it. Each remaining host comes with suggested wording for the approach, and the list is re-derived every check, so you can watch it shrink.

What does it mean when Google shows no AI Overview for a question?

It means the engine was read and had nothing to show — Google's results page carried no AI Overview for that query. It's recorded as its own state, counted separately, and kept out of every denominator, because it is neither an answer nor a failure. The important part is what it is not: it is never read as "you weren't named". That would turn an absent answer into a fact about your brand, which the data cannot support. The same care applies in the other direction — a question we simply couldn't read is our problem and enters no count either. Answered, no answer shown, and not measured stay three distinct states everywhere in the dimension.

Last updated on