Skip to content

AI Competitor Lists: A Definition of Your Market, Not a Ranking

August 14, 2026
15 min read
AI Competitor Lists: A Definition of Your Market, Not a Ranking

The fastest way to build a competitor analysis today is to ask an AI which tools lead your category. Ask ChatGPT, Claude and Gemini the same question and you get three lists. The obvious thing to compare is the ranking - who came first, and whether you are on it at all.

That is the wrong reading, and it fails in two stages. Across 50 software categories pulled from a public AI-visibility observatory, a large share of what gets reported as "the engines disagree about who leads" turns out to be arithmetic: the engines write lists of different lengths, so comparing them whole partly measures the length. Strip that out and a different disagreement is left standing - one we could not find measured anywhere across three research passes. The three lists do not rank one market. They describe three different ones.

What follows is the finding, the correction that makes it visible, and an honest account of what a category list is evidence of. Every figure comes from a public API under an open licence, and the limits are stated in the body rather than a footnote.

Key Takeaways
  • Asked who leads competitive intelligence software, the three engines returned structurally different lists. Compared at the same length, ChatGPT's exclusive picks were advertising and social listening tools, Claude's were dedicated vendors, and Gemini's was a review site - and past rank ten Gemini's list ran on to Crunchbase, LinkedIn and Statista.
  • Build a competitor set from ChatGPT's ten and four of the names sell advertising tools, while AlphaSense and Contify - two vendors that do sell competitive intelligence - are missing entirely.
  • Compared list-for-list at the same length, seventeen companies are in play and only four appear on all three: Crayon, Kompyte, SimilarWeb and Semrush. Much of the reported "engine disagreement" is arithmetic - correcting for list length across 50 categories moved the share of companies named by only one engine from 55% to 48%, and reversed which engine looked most idiosyncratic.
  • ChatGPT never wrote Klue into its own list of ten, yet discussed it on sight in a published transcript and then endorsed a shortlist containing it. Across 49 categories, 38% of agreed companies were absent from ChatGPT's own five - the highest of the three.
  • In a skincare experiment, well-known brands won 100% of model recommendations when product specifications were identical. We could not find an equivalent experiment on B2B software.
  • Do this: cut every engine's answer to the same length, sort the names into buckets by what kind of company they are, and read the shape of the list before you use it as a competitor set.

What do three AI engines think is in your category?

Not the same thing. Ask ChatGPT, Claude and Gemini who leads competitive intelligence software and you do not get three rankings of one market. You get three different answers about what the market is.

Start with the thing that makes this possible: none of the three is reading out a stored ranking. Each generates its answer in response to the question, which is why the same question asked twice can return a different list - an effect measured further down.

The lists here come from Respectarium, a public AI-visibility observatory that puts one category question to ChatGPT, Claude and Gemini every week and publishes the answers under an open CC BY 4.0 licence. Figures are from the cycle of 9 August 2026, retrieved 13 August 2026. Full method at the end.

The three lists are also not the same length. ChatGPT wrote ten names, Gemini nineteen, Claude twenty. So before comparing anything, every list is cut to ten - the length of the shortest. That cut turns out to matter enormously, and the section on list length below shows why.

Cut to ten, the three engines name seventeen companies between them. Four of those seventeen appear on all three lists: Crayon, Kompyte, SimilarWeb and Semrush. The other thirteen are one or two engines' opinion.

Three column comparison of ChatGPT, Claude and Gemini competitive intelligence lists showing only four companies shared across all three

The interesting part is what each engine named that neither of the others did. We sorted every company into one of five buckets - dedicated competitor, adjacent tool, directory or marketplace, review site, data provider - and the exclusive picks fall out like this.

EngineNamed only by this engineBucket
ChatGPTAdbeat, BuzzSumo, RivalIQ, BrandwatchAdjacent tools - advertising and social listening
ClaudeContify, AlphaSense, VisualpingDedicated competitors
GeminiTrustRadiusReview site

Let the longer lists run past ten and Claude and Gemini diverge further in the same directions. Claude carries on through more dedicated vendors. Gemini carries on through Crunchbase, ZoomInfo, Product Hunt, LinkedIn and Statista - a professional network and a statistics portal, returned in answer to a question about software tools.

Sorting companies into kinds of product is our judgement, not the source data's, which records only who was named and at what rank. But grouped that way the shape is plain: across their full lists, Claude's answer is mostly dedicated vendors, Gemini's is mostly directories and data platforms, and ChatGPT's leans on advertising tools.

The named companies grouped by what kind of product they are, showing each engine draws the category boundary differently

What a bad AI competitor list does to your benchmarking

Ask an assistant who your competitors are and you get an AI competitor list. Build your competitor set from it and you inherit that engine's definition of your market.

Say you sell competitive intelligence software and you take ChatGPT's ten. Four of those names - Adbeat, BuzzSumo, RivalIQ and Brandwatch - are advertising and social listening tools. None of them sells what you sell. Meanwhile AlphaSense and Contify, two vendors that do, are missing entirely: only Claude named them.

You would now be benchmarking your pricing against ad-tech, reading your feature gaps off social listening tools, and writing battlecards for two competitors you never knew you had. Take Gemini's list instead and the same competitor analysis is built substantially out of directories and data providers. Neither list answers "who are my competitors" the way you would answer it yourself.

That failure is invisible if you read the list as a ranking. It is obvious the moment you read it as a definition.

Across three research passes we could not find anyone who has classified AI category answers by what kind of company appears in them. Studies counting whether a brand was named are common; we found none asking whether the named thing belonged in the category at all.

Do the engines really disagree about the boundary, or are we reading it in?

They argue it out loud, by name. Alongside the weekly rankings, the observatory publishes a transcript per category in which the three engines' picks are argued out and an agreed five emerges.

Each turn is attributed to a named engine and timestamped. The competitive intelligence exchange ran eight turns in 41 seconds, retrieved 13 August 2026. How the exchange is produced is not documented publicly, so it is worth what is visible on the page and no more.

What they argue about is not who is best. It is what counts:

"Gemini's pick of Ahrefs is off-target. Ahrefs is an SEO tool, not a competitive intelligence platform." - Claude

"ChatGPT's inclusion of Owler is questionable - it's more of a basic company news aggregator than a true CI platform." - Claude

Later in the exchange Gemini gives way: "I concede on Ahrefs; while powerful for SEO, it doesn't fit the core CI platform definition as Claude and chatgpt correctly argued." Both contested names are dropped, and the three settle on a five-company shortlist: Crayon, Klue, Kompyte, SimilarWeb and Semrush.

An engine's list is not an inventory of what it knows

Klue, a dedicated competitive intelligence vendor, makes the point. Claude ranks it first and Gemini second. ChatGPT does not name it anywhere in its own ten.

Yet in the transcript ChatGPT discusses Klue on sight, calling it "an interesting pick, favored by Claude and Gemini" and arguing it is "less versatile than Kompyte." It can evaluate the company perfectly well once the name is in front of it. Then it signs off on a shortlist with Klue on it.

That rules out one explanation - the engine does not fail to recognise Klue - without settling what did happen. It may not have retrieved the name when writing its own list, or it may have retrieved and dropped it. We cannot tell from the outside, and no published work lets anyone read an endorsement-after-challenge as proof the model knew all along.

That is not a one-off. Compare each engine's own five best picks against the five the three of them settled on, and across the 49 categories that reached agreement, 38% of the agreed companies were absent from ChatGPT's own five. That is the highest of the three, against 23% for Claude and 20% for Gemini - and higher means an engine endorsing more companies it had not named itself.

Bar chart showing 38 percent of agreed companies absent from ChatGPT's own top five versus 23 percent for Claude and 20 percent for Gemini

Be careful about what a concession proves. Debate or Vote, presented at NeurIPS in 2025, a leading machine-learning research conference, showed that rounds of debate carry no built-in pull toward the right answer, and that most of the benefit people credit to models arguing comes from simple majority voting. A model that concedes may be revealing knowledge. It may just be agreeable.

The recognition is the durable part. Whatever the concession is worth, an unprompted list is a sample of what an engine wrote down, not an inventory of what it can recognise. There is one transcript per category and no history behind it, so this is 49 categories each measured once - a snapshot, never a trend.

Isn't this just the engine disagreement everyone already reports?

No, and most of that reported disagreement is arithmetic. Published estimates of how much the major engines overlap when naming category leaders run from 15% to 83%. That spread is too wide to be about the engines. It tracks a choice each research team made about list length before collecting anything.

StudyHow the lists were comparedAgreement reported
SandsDX, an AI-visibility research firmEvery model capped at ten names; top-five shortlists compared83% (strongest model pair)
BrightEdge, an SEO software vendorWording not disclosed; top 100 names per engine, in aggregate36-55%
SEOForge, a search agencyNo length limit; every name compared29.5%
Parse.gl, an AI-search analytics firmNo length limit; every name compared25%
SMA Marketing, a search agencyNo length limit; top 3 pulled out afterwards15%

Only the top row told every model to return the same number of names before comparing them. All five are vendor or agency research, none of it disinterested - and neither are we.

These five figures are not comparable with each other either. One scored the top five of a capped ten-name answer, others scored whole lists, a top hundred, or a top three. Agreement measured on lists of different lengths, lined up in a column as though the column meant something. That is the error the published record commits before any re-analysis begins.

Why length alone moves the number

The usual score divides the names two engines share by the total number of distinct companies across both lists.

Say seven of ChatGPT's ten names also appear on Claude's twenty. The two lists mention twenty-three distinct companies between them: ten from ChatGPT, twenty from Claude, minus the seven counted twice. Seven of twenty-three is 30%. Seven out of ChatGPT's own ten is 70%. The shared names never changed. Only the yardstick did.

Push it to the limit. If all ten of ChatGPT's names appeared on Claude's list - complete agreement, nothing in dispute - the usual score still reaches only 50%. Perfect agreement cannot score better than half.

You never need a rule of thumb when the ceiling can be calculated. Across the 50 categories, the highest score ChatGPT and Claude could possibly have reached against each other is 54%. They scored 34% - about 63% of what was achievable, not the collapse it looks like against a scale that appears to run to 100. Scored against the shorter list instead, the same fact reads differently: 71 of every 100 tools ChatGPT named also appeared on Claude's list.

What changed when we recounted

We merged the fourteen cases where one company appeared under two domains, leaving 1,256 entries - one company, named by one engine, in one category - then cut all three engines to their top ten and recalculated across all 50 categories.

MeasureAs normally reportedAfter cutting every list to ten
Entries being compared1,256810
Share named by only one engine55%48%
Share named by all three24%31%
Total named by ChatGPT / Claude / Gemini537 / 989 / 592499 / 497 / 479

The first row is the one to watch, and it is the reason this table needs reading carefully. Cutting every list to ten does not just reweight the corpus, it shrinks it: the entries ranked eleventh and below stop existing. So the two columns are shares of different totals, which is exactly the trap described above. We print both denominators rather than let the percentages float.

Look at the per-engine counts. Claude appeared to name almost twice as many companies as ChatGPT. Compare the same ten positions and the gap vanishes: all three name roughly 500.

The engine that swapped places

Then the direction of the finding reverses. Count the companies each engine named that no other engine did - a higher number meaning more names nobody else picked.

On raw counts Claude is the obvious oddity at 457, against ChatGPT's 126. Equalise the list lengths and ChatGPT becomes the outlier at 173, against Claude's 125.

Slope chart showing Claude's single-engine company count falling from 457 to 125 while ChatGPT's rises from 126 to 173 once list lengths are equalised

Same data. Same week. Opposite conclusion. The only thing that changed was comparing like with like.

Credit where it belongs: SandsDX found this effect before we did and published a correction when one of their own reported findings failed a list-length test. Theirs is the most transparent methodology we found, and it is also the study reporting the highest agreement. Those two facts are related.

The strongest objection. A longer list might be a real behavioural difference. If one engine habitually pads its answers with marginal tools, penalising it is arguably correct rather than a flaw. That is a fair thing to measure - but it is a question about list length, and you answer it by reporting list length. Folded into a figure labelled "agreement", it invites readers to conclude the engines named different companies, which is mostly not what happened.

One finding is immune to the list-length problem specifically, because a number one is a number one whatever the list length. It is not immune to the number of engines you compare: agreement on the top spot falls as you add models, and one vendor study across eight models put unanimous agreement on the leader at 4%. Across the 50 categories all three engines picked the same leader in 24. In 21 more, two agreed and one differed. In only 5 did all three name something different. Parse.gl found the same shape at larger scale: across 1,655 categories, ChatGPT and Google's AI picked the same number one 93% of the time, while their full lists - uncapped, and so subject to the same length effect - overlapped 25%.

What does presence on a category list measure?

Accumulated public evidence about a company, rather than whether the company belongs in the category.

The cleanest demonstration comes from outside software. In Incumbent Advantage, published June 2026, the researchers Chu and Hou showed three commercial models sets of skincare products whose specifications were identical, varying only the brand. Well-known brands took 100% of the recommendations. That dominance was fragile: it broke once a rival gained less than a tenth of a star.

When the specifications are level, the model falls back on name recognition. Whether that carries from skincare to B2B software is untested, and we could not find anyone who has run the equivalent experiment on software categories.

Very little of this is settled. A study of 1,094 categories published in July 2026 by Semrush, which sells AI-visibility tooling, found a clear leading brand in only 15.2% of them in ChatGPT. More than half were unsettled entirely.

We can speak to the other side of it directly. CompetLab - we build competitive intelligence software, and this is our blog - is named in none of the 50 categories this observatory covers, re-checked against the live API on 13 August 2026. Neither is Otterly, nor ChampSignal, two of the competitors we name on our own site. Three tools built for this market, none of us named anywhere in it. That is what you would expect if these lists mostly reward companies with a long public record, and a separate self-test we ran on 12 August 2026 - nine prompts put to all three engines, twice each, for 54 queries - returned the same answer: zero mentions, on every engine.

How do you analyse your competitors' visibility in LLMs?

Read composition rather than position, at matched length, more than once. The four steps below turn a raw AI competitor list into something you can benchmark against.

  1. Cut every answer to the same length before comparing. If one engine gives you ten names and another twenty, comparing them directly invents disagreement that is not there. Use the shortest list as the cut.
  2. Tally what kind of company came back, not just whether you appeared. Put every name into one of five buckets: dedicated competitor, adjacent tool, directory or marketplace, review site, data provider. A spreadsheet with one row per name, one column per engine and a bucket column is enough. Then read the shape. If an engine answers your category mostly with directories, its answer is drawn from a different market than the one you sell into, and any competitor analysis you take from it will point your pricing and positioning at the wrong companies. This is the step almost nobody does, and it is where the value is.
  3. Run the same prompt again. One answer is one roll of the dice, and repetition alone moves it. MaxAEO, an AI-visibility vendor, found that asking an engine the identical question twice changes at least one named brand 17% of the time, across 300 repeat checks.
  4. Act on the definition, not the rank. If two engines describe your market the way you do and one does not, the question is not how to climb that one's list. It is whether your own site says plainly what category you are in, and which definition your buyers are being handed.

Track how often you appear across many runs rather than your position in any single answer. That is the reasoning we set out in our guide to measuring AI visibility, and it is the metric this whole article points at.

How long is a reading good for?

The top of a category and the rest of it move at completely different speeds.

We walked the observatory's weekly archive across 30 categories, giving 507 week-to-week comparisons between 13 April and 9 August 2026. The leading company changed in only 25 of those 507 comparisons - under 5% of weeks. The composition of the top five changed in 42% of them. A company present in both weeks moved up or down about two positions.

Read those two numbers differently, because only one of them clears the noise. Each weekly observation is a single run, and the section above put run-to-run turnover at 17% for an identical repeated prompt. A 42% weekly change in the top five therefore cannot be separated from ordinary sampling wobble - it bounds the churn from above rather than measuring it. The sub-5% figure for the leader is the durable one: it holds despite that noise, which is what makes near-total stability at the top the real finding.

Two more caveats. This tracks Respectarium's single merged ranking - one list built by blending all three engines' answers - not churn inside any single engine's list. And categories differ sharply. Competitive intelligence, the category this article opened with, barely moved: its top five changed in 2 of its 17 week-to-week comparisons, against 42% across all 30, and its leader never changed once.

The practical reading: the top of a category is stable enough to treat as a fact and worth re-checking monthly. Everything below it is too noisy to read as a position at all, at any cadence.

What the data does not show

Nothing we could find shows that absence costs you money. There is no controlled study connecting presence in an AI recommendation set to pipeline or revenue - only vendor case studies with self-attributed numbers and no control group. The widely quoted figures for AI-cited vendors, 2.3x demo rates and 34% shorter sales cycles, come from a vendor in the space and are directional at best. What buyer research does support is narrower: people are asking assistants for vendors, and zero visibility means zero chance of being in the answer. Between those two facts is a gap we found no measurement of.

The recount has limits too

It rests on one answer per engine. The observatory queries each engine once per category per week, so every figure above rests on one run - the exact weakness step 3 warns about, applied to us. What holds up better is that we collected no new data: we recounted published answers with a fairer yardstick.

What the finding does not settle

We do not know which question buyers actually ask. Real prompts are often constrained - "three CRM tools for a hospital that work on iPads" - but there is also credible evidence that broad category questions are a common opening move when a buyer does not know where to start. We could not find the fraction published, or any research comparing the vendor sets the two styles produce. That cuts both ways: absence from a broad list is neither proven harmless nor proven fatal.

The three lists may differ because of plumbing, not judgement. We read the composition gap as three definitions of a market. A competing explanation is that the engines run different retrieval stacks, and a benchmark of 3,996 cited sources across twelve systems found that pairing a weaker model with stronger search beat a more capable model with weaker search on source quality. Some stacks also carry licensed data partners, which would put a Statista or a Crunchbase in reach of one engine and not another. We cannot separate the two from outside, and either way the consequence for your competitor set is the same.

Absence may measure age rather than fit. These lists favour companies with years of reviews and coverage behind them. A fair reader can conclude they mostly describe incumbency, and on this evidence that reader is right.

One further limit arrives on a date. From 15 September 2026 Cloudflare blocks training and agent crawlers by default on ad-carrying pages, for domains onboarding after that date. Where it applies, a company missing from an AI answer may be reflecting its own crawler policy rather than its market position - and from outside the two look identical.

None of this touches the finding that the three lists describe different markets. Whether or not absence costs you revenue, a list built from a different definition of your market is the wrong list to benchmark against.

The engines do behave differently, and that is worth tracking - we found Claude and GPT rewarding the opposite things. But those differences change how an engine describes a company it already knows; they will not get an unknown company named at all. Most advice runs the two together.

Methodology. Rankings and debate transcripts come from Respectarium's public API under a CC BY 4.0 open licence. Cycle of 9 August 2026, retrieved 13 August 2026: 50 categories, 1,270 raw entries, 1,256 after merging fourteen duplicate-domain pairs. Its public methodology page names the models queried - gpt-4o, claude-sonnet-4.5 and gemini-2.5-flash - and states that one prompt per category goes to all three engines with no starting list. List lengths are a stable habit rather than a cap: ChatGPT returned exactly ten names in 43 of the 50 categories, Claude exactly twenty in 42, Gemini between seven and twenty. Matched-length figures cut every engine to its top ten - the depth ChatGPT actually reaches in most categories, and the shortest of the three lists in the category examined here. One run per engine per category per week, so run-to-run variance cannot be measured from this dataset, and the debate corpus holds one snapshot per category with no archive behind it. The week-to-week figures walk the 18-cycle archive (13 April to 9 August 2026) for the first 30 categories, giving 507 consecutive-week comparisons; they track the single merged ranking, whose published weighting is 40% average position across the three engines, 35% how many of the three named the company, 15% whether the mention was a recommendation or an aside, and 10% the engine's own stated confidence. Classification of companies into the five buckets is ours, single-coded by one reviewer, and is not in the source data. CompetLab is not involved in producing the ranking data.

Want to see which companies AI puts you beside?

CompetLab tracks how ChatGPT, Claude and Gemini rank and describe your brand, and which companies they put you beside.

14-day free trial. No credit card. Start here.

Frequently Asked Questions

Can I build a competitor analysis from an AI's category list?

Not without checking what is on it first. In the category we examined, compared at matched length, one engine's exclusive picks were advertising and social listening tools, another's were dedicated vendors, and the third's was a review site - with directories like Crunchbase and LinkedIn appearing further down its longer list. Take the first list and you would benchmark pricing against ad-tech while missing two real competitors entirely. Sort every name into buckets - dedicated competitor, adjacent tool, directory, review site, data provider - and read the shape before you use it.

Why do the three AI engines name different companies for the same category?

Two reasons, and separating them matters. Part is arithmetic: the engines return lists of different lengths, so comparing a ten-item list against a twenty-item one manufactures divergence that is not there. Cutting all three to the same length across 50 categories moved the share of companies named by only one engine from 55% to 48%. What remains after that correction is a genuine disagreement about what the category contains, not merely about who leads it.

Why doesn't ChatGPT mention my company?

Most likely because there is not enough public evidence tying your company to the category. Note this is not the same as the model never seeing you: in a June 2026 experiment on skincare products, the researchers Chu and Hou found that when three commercial models were shown products with identical specifications, well-known brands took 100% of recommendations, and that lead broke once a rival gained less than a tenth of a star. We could not find an equivalent experiment on B2B software.

Does being absent from an AI category list mean buyers cannot find me?

Not necessarily, and across three research passes we could not find anyone who has measured it. There is no controlled study connecting presence in an AI recommendation set to pipeline or revenue - what exists is vendor case studies with self-attributed numbers and no control group. Absence from a broad category prompt also does not mean absence from the narrower, more specific questions real buyers ask. Treat a missing listing as a visibility risk, not a proven revenue loss.

How often do AI category rankings actually change?

The leader barely moves; everything under it does. Across 30 software categories and 507 week-to-week comparisons in a public weekly archive, the top-ranked company changed in under 5% of weeks, while the composition of the top five changed in 42%. A company present in both weeks moved about two places on average. Two caveats: this is a combined three-engine leaderboard from Respectarium, not any single engine's list, and each week is one run, so the 42% cannot be separated from ordinary run-to-run noise. Treat the top as stable and everything below it as unreadable.

How many times should I run the same prompt before trusting the answer?

More than once, because repetition alone moves the result. MaxAEO, an AI-visibility vendor, found that asking an engine the identical question twice changes at least one named brand 17% of the time, across 300 repeat checks. Track how often you appear across many runs rather than your position in any single answer, and treat one response as one roll of the dice.

Watch your whole market, automatically

Track competitors' pricing, positioning, content and AI visibility — with a monthly Strategic Briefing on what changed and what to do.

Share this article