Methodology
This page documents how the benchmark indicators are built, including the data exclusion criteria. Every figure is read from the database when the page loads.
What is measured
Structured sets of questions are sent to 3 language models (ChatGPT, Gemini and Claude) about 15 industries across 20 markets. Each response is checked for which brands are mentioned, in what order, whether an explicit recommendation is made, and the tone of the context each mention appears in.
Queries run against each model family's current API version. The exact identifier of each model is recorded with every snapshot: ChatGPT (gpt-4o-mini-2024-07-18), Gemini (gemini-3.1-flash-lite), Claude (claude-haiku-4-5-20251001) (from the 31 August 2026 snapshot).
Cumulative total: 277,815 responses analysed since .
Mention detection and exclusion criteria
Brands are extracted from the text of each response. This process produces two classes of false positives, and both are filtered before anything is published.
Generic terms. Extraction can capture section headings and ordinary nouns as if they were brand names. A curated brand registry is maintained, and only entries listed in it are published. The 100 highest frequency entries were reviewed manually.
Substring collisions. A short brand name can appear contained inside an unrelated word. Each mention is validated by requiring the name to appear as a whole word in its context. Without this validation, short brand names would accumulate mentions that do not refer to them.
Figures from the 29 August 2026 run, over 1,672,060 mentions processed. Recomputed weekly.
- Counted
- 62.9%
- Outside the brand registry
- 31.5%
- Excluded, substring collision
- 5.6%
Of the subset of mentions that do match a registry brand, 8.1% turned out to be a false positive and was excluded.
When we publish a view
A market crossed with an industry, for example Insurance in Uruguay, is a view. Not every one is published: below a certain volume, the ranking would be more sampling noise than a real reading of the market.
- A view publishes if at least one of its two most recent snapshots reaches 90 verified mentions or more, and 8 brands or more with 3 mentions or more each.
- The better of the two snapshots is evaluated, not only the latest, so a view does not appear and disappear from one week to the next over a sampling swing that crosses the threshold in either direction.
- A view that does not clear both requirements in either snapshot does not render: it shows the same state as a combination that has not had a snapshot yet.
Inside a published view, no row with fewer than 3 mentions appears in the table, the share of voice donut or the per model podium. A view with few brands above that floor publishes fewer than ten rows rather than filling the rest with brands that do not clear it.
Criteria for publishing a movement
Each snapshot rests on a limited number of responses. At that volume, a brand with few mentions can shift several positions from a minimal variation. Publishing that shift as a trend would amount to reporting sampling noise.
For that reason minimum thresholds apply:
- Row level movement indicators require at least 5 mentions in both snapshots compared.
- The biggest movers section requires at least 8 mentions in both snapshots. The threshold is higher because that section asserts a movement was among the largest of the period.
- The minimum is required in both snapshots, so no brand is credited with a large rise on the strength of a single favourable snapshot.
The absence of an indicator does not mean the absence of change. It means the movement cannot be measured with statistical confidence. Those cases are marked "no data" or "new", never a zero. They are different claims and are kept separate.
Measurement cadence
Measurement is grouped into weekly snapshots. Meaningful changes in model answers happen when a version updates or when the content it draws on changes, processes whose time scale is weeks, not days. A daily frequency would multiply the volume of queries without adding signal.
Each market is surveyed every one to three weeks, so some market and industry combinations have weeks with no measurement. Those gaps are not interpolated: the trend line breaks where there was no measurement, since drawing a line across an unmeasured period would assert a trend that was not observed.
Trend lines are shown from 3 snapshots onward. Below that minimum, the page shows that the series is still being built, and since when.
Date of each snapshot
27 snapshots on record, newest first.
Known limitations
- Tone is determined through automated analysis of the context around each mention, not through human review. It is an indicator of general direction and should not be used for fine comparisons between brands.
- Providers update their models without advance notice. A movement in the ranking can originate in the market or in a model change, and the two causes are not always distinguishable from outside.
- The Brazilian market is queried in English. Its ranking reflects the models' answers in that language about the Brazilian market.
- A new brand does not appear in the rankings until it accumulates the minimum volume required to enter the registry. This criterion prioritizes excluding noise over immediate coverage.