Understand · AI visibility figures
What are AI visibility figures really worth? What the market’s tools, and ours, actually measure
An AI visibility figure never describes your company: it describes a measurement procedure, and it is only as good as that procedure. Two honest studies can therefore report very different results without either being wrong, simply because they did not measure the same thing. This page explains how to read these figures, including ours, and which questions to ask before paying for one of them.
We are writing this page because we sell measurement. A reader is therefore entitled to suspect us, and the only acceptable response is to publish the method, the limitations and the sources, then let the reader judge. Every address cited here was opened and checked while this page was being written.
What exactly does an AI visibility tool measure?
When a tool tells you that your “AI visibility” is 42, it has made at least four decisions on your behalf, and none of the four is neutral.
The first decision concerns the surface being tested. An answer engine can be reached through several doors: the application your customer uses in a browser, an application programming interface, which is the technical channel through which software talks to the engine without using that screen, or the command-line tools distributed by the providers themselves. These doors do not produce the same answers, and the next section shows how wide the gap can be.
The second decision concerns the questions. Thirty buying questions do not measure the same thing as thirty advisory questions. During our reading on the night of 14 to 15 August 2026, the ChatGPT application triggered a web search for only 11 of the 30 questions asked: for the other 19, it answered from memory, without citing anyone, neither us nor a competitor. A tool that mainly chooses questions for which the engine does not search will mainly measure zeros, while a tool that mainly chooses buying questions will mainly measure citations. Both will be right, and neither will have informed you.
The third decision concerns the number of trials. An AI answer is a draw, not a ranking: asking the same question twice can produce two different answers. A figure obtained in a single pass is therefore not a measurement. It is a sample of size one.
The fourth decision concerns what is accepted as a citation. Being used as a source, being named in passing in the text and being recommended to a buyer are three separate situations, which rule 05 of our measurement charter requires us to separate. A tool that adds them together mechanically produces a more flattering figure than a tool that distinguishes between them.
These four decisions can be read in the denominator, which is why we always publish ours. In our public Real French diagnostic, the same result is written in two ways: 6 answers out of 270, and 4 question-and-engine combinations out of 90. The two fractions describe the same measurement. They do not tell the same story, and a single figure without its denominator cannot be verified.
An English-language study concludes that the API and the application are two different engines. What is our answer?
On 13 August 2026, the Paris company H Company published a study by Antonio Ventura entitled “Learning From the User: GEO Using Computer-Use Agents vs. API limitations”. Across 30 real buying questions and several passes per question, it compares the results returned by OpenAI’s application programming interface with web search against what the ChatGPT application shows to a user.
Its findings deserve to be read in full. About 18% of the products highlighted by the application also appear in the application programming interface’s answers, or fewer than one in five. The overlap between cited links falls to about 2.5%, even though the application programming interface cites roughly twice as many links and suggests more products per question, 6.2 compared with 3.6 on average. Nor are the sources of the same kind: on the application programming interface side, almost 78% of the evidence comes from third-party editorial content, while the application draws more heavily from brand websites and community discussions. The authors sum up their conclusion in a sentence that we reproduce in their own language: “API and UI behaved like two different web research engines, with different retrieval behavior and meaningfully different evidence pools”.
On its own ground, this study is right, and we have no desire to soften it. It establishes a fact that many buyers do not know: connecting a tool to an application programming interface and presenting the result as “what ChatGPT shows your customers” is an equivalence that nobody had verified, and that does not hold for the buying questions tested in this study. A reader who is about to subscribe to a tracking tool should read it before signing.
Here is what it does not say, and this must be stated just as precisely. It compares two surfaces: the application programming interface and the application. It does not compare the command-line tools distributed by the providers themselves, which form a third door and are the route taken by our own readings. No provider publicly documents whether its command-line tool’s web search shares its application’s infrastructure. We therefore know no more than anyone else, and we will not claim that this finding does not concern us.
What we did was measure the result instead of debating it. On the night of 14 to 15 August 2026, we asked the 30 frozen questions from the Real French diagnostic on both sides, during the same night: a human in the ChatGPT application, and our measurement instrument. The two surfaces returned the same presence verdict on 30 questions out of 30. The protocol, the single divergence, the limitations and the evidence are published in our correlation study, and the reading will be repeated every quarter.
Both findings are true at the same time, and the next section explains why. We add a precaution that works both ways: H Company sells an agent-based measurement platform, and we sell measurement. Neither company is a neutral third party, which is precisely why each study must be judged on its published method rather than on its conclusion.
How can two honest measurements produce thirty out of thirty and eighteen per cent?
Because they do not count the same thing. One difference accounts for the entire gap, and it can be understood without advanced statistics.
H Company measures the overlap between two lists. Take the list of products recommended by the application, take the list from the application programming interface, and see what proportion appears in both. It is a fine-grained measurement, and it is severe by design: the overlap collapses as soon as the second list suggests different brands or different comparison articles. The authors also used a measurement of this kind, Jaccard similarity, to assess the stability of their own passages.
Our reading measures something else: a presence verdict, for one named brand, question by question. Is Real French cited as a source, mentioned in the text, or absent? The answer is one word, and the two surfaces returned the same word thirty times in succession. Two engines can cite entirely different websites and agree that your company does not appear in either set.
Take a question for which the application cites eight sources and our instrument cites eight other sources, with none in common. The overlap is zero, while the presence verdict is identical on both sides because the brand being sought is absent from all sixteen sources. The same reading therefore produces 0% and 100% agreement depending on the question asked of the data. A reader who translates “18% overlap” as “this instrument is 18% accurate” has changed the subject along the way.
One objection remains, and it is the right one because we make it ourselves. A high agreement rate for a rare event is easy to obtain. If a brand is absent from almost every answer, two instruments that blindly answer “absent” will agree almost 100% of the time. In 1960, Jacob Cohen made precisely this criticism of a simple agreement rate, which he argued could not account for agreement obtained by chance, and for this reason he proposed a corrected coefficient. Our 30 out of 30 firmly shows that our instrument does not invent visibility. It does not yet show that it sees a citation when a citation really exists.
What has our own study not yet demonstrated?
Five limitations are written on the study page, and we repeat them here because they matter more than the result.
- The brand tested is almost always absent. The reading therefore tests mainly the instrument’s specificity, meaning its ability not to report a presence that does not exist. It is a poor test of its sensitivity, meaning its ability to detect a brand that is genuinely cited.
- The application was tested only once per question. An application answer is a distribution, with its memory, location and the tests in progress at the provider. We sampled one point per question, on one account and during one night.
- The control covers only one family of engines. It covers the OpenAI surface. Our instruments for Claude and Gemini still need to be controlled in the same way, and we make no claim about them until that work has been done.
- The instrument has changed since then. Since 16 August 2026, our probe has queried the model served to paying ChatGPT subscribers. The published control compares the previous configuration with the application: it remains true for that configuration, and no longer describes the one now in service.
- The agreement coefficient cannot yet be calculated. Because every verdict falls into the same class, a coefficient such as Cohen’s is not even defined. It will become calculable with a varied panel, which is the purpose of the November 2026 reading. That reading will cover brands that are always cited, sometimes cited and never cited, with quantified sensitivity and specificity.
One final limitation goes beyond our instrument and applies to all of them. On 7 August 2026, for the question “one-on-one French immersion homestay in France”, Real French was cited in all three passes, in source positions 3, 6 and 7. One week later, during the night of 14 to 15 August, none of the tested surfaces cited it. A citation can therefore appear and disappear within seven days, without anything having changed on the website. No tool, whatever its price, turns a snapshot into something permanently acquired: only repeated, dated measurement informs you.
We publish these limitations because a figure whose limitations are unknown is worth nothing. We apply the same rule to our own name: in the reading of 11 August 2026, HENUSSE was not retained in any of the 30 answers to questions in which a buyer is looking for a provider, and that zero is published as it stands in our Barometer.
What questions should you ask a tool provider before subscribing?
Here are the seven questions we would ask in your position, of any AI visibility tracking tool provider. They are useful to the reader even if they buy nothing from us, and we recommend that you use them against us too: if our answers seem weak, that is where the conversation should begin.
- Which surface do you test, and do you state it on every reading? The application, an application programming interface, a command-line tool: the answer must be a name, not the word “ChatGPT”. This is the question that the H Company study makes unavoidable.
- Which exact model answered, and on what date? Models change in a matter of days, and a reading without its model and date cannot be repeated. A serious provider will also tell you what it did on the day the model changed.
- How many times was the same question asked? You should see a fraction, “cited 2 times out of 3”, not a binary status. A single pass measures nothing, because the next answer may be different.
- What do you call a citation? Ask for a distinction between being cited as a source, being mentioned in the text and being recommended to a buyer. A tool that refuses to separate the three is selling a higher figure, not more accurate information.
- Can I read the raw answers and repeat the test? A provider that keeps the complete answers, the pages used, the time and the time zone allows you to recount after it. A dashboard that shows only charts cannot be verified.
- Have you published a control comparing your tool with the application, including its limitations? An agreement rate alone is not enough: ask which brands it was established on, and require sensitivity as well as specificity. Perfect agreement obtained on brands that are absent everywhere is easy to demonstrate, and we are the first people to whom that criticism applies.
- Were your questions frozen before the measurement, and are the zeros published? Choosing the questions after seeing the answers can produce any result. Publishing the zeros is the best sign that a provider is not selecting its figures.
One point is worth knowing because it shows the maturity of the field. In a mature field, an instrument that returns a verdict is described according to a published list: in medicine, the report of a diagnostic accuracy study follows the STARD checklist, thirty published items that require the authors to name the test being evaluated, the standard against which it is compared, the population tested, the handling of missing results and the study’s limitations. Its very first item requires the report to announce at least one measure of accuracy, such as sensitivity or specificity. Nothing of this kind yet exists for AI visibility. In the meantime, these seven questions serve in its place, and the only thing a buyer can require today is that every figure comes with its surface, date, fraction and definition.
These questions concern the measurement tool. The questions to ask the provider working with you are different, and we publish ours, with our answers, on the page about questions to ask before signing.
Frequently asked questions
Does the H Company study invalidate your figures?
No, because it does not measure the same thing. It compares the overlap between two lists of products and links, while our reading compares a presence verdict for one named brand. Both findings hold together, and its result imposes an obligation on us that we accept: to continue controlling our instrument against the application, on published dates.
Should I prefer a tool that measures directly in the application?
That surface has one obvious quality, because it is the one your customer uses, and an equally real measurement flaw: an application answer depends on the account asking the question, its memory, its history and its location, which makes it difficult to repeat under identical conditions. The right question is therefore not which side to choose, but what each provider has measured about the gap between its surface and the one your customer uses.
Is 100% agreement enough to validate an instrument?
No, and this is the main limitation of our own study. High agreement obtained on a brand that is almost always absent proves that the instrument invents nothing, not that it detects a citation when one exists. That requires a varied panel of brands, which is what we will publish in November 2026.
Do these figures apply to my company?
None of the figures cited on this page applies to your company, and that includes H Company’s figures. Each concerns a precise set of questions, a precise surface and a precise date. The only figure that concerns you is the one obtained from your buying questions, frozen before measurement.
What will you do if your November reading proves you wrong?
We will publish it in figures on the same page, and we will adjust the instrument. The rule is written in our measurement charter: the questions are frozen before the test, every fraction carries its date, and a zero is written as a zero.
Sources
- H Company, Antonio Ventura, “Learning From the User: GEO Using Computer-Use Agents vs. API limitations”: https://hcompany.ai/geo-using-computer-use-agents-vs-api-limitations. Contains the 30 buying questions, the multiple passes, the overlap of about 18% for products and 2.5% for links, the 6.2 products compared with 3.6, the 78% share of editorial sources, the use of Jaccard similarity and the concluding sentence quoted in the text. Accessed and checked on 16 August 2026. The page header shows a February 2026 date, although the page went online on 13 August 2026: it does not appear in the provider’s blog index, which was built on that same 13 August a few minutes earlier. We use the publication date and disclose the discrepancy instead of concealing it.
- HENUSSE, correlation study between the ChatGPT application and our instrument: https://henusse.com/en/ai-search/diagnostic/correlation-study/. Contains the 30 out of 30 from the reading on the night of 14 to 15 August 2026, the 11 questions out of 30 for which the application triggered a search, the limitations repeated here, the instrument change of 16 August 2026 and the case of the citation that disappeared between 7 and 14 August 2026.
- HENUSSE, Honest Measurement Charter: https://henusse.com/en/ai-search/measurement-charter/. Contains the nine rules cited, including the distinction between cited, mentioned and recommended, the repetition of questions, the required date and the publication of zeros.
- HENUSSE, public Real French diagnostic: https://henusse.com/en/ai-search/diagnostic/example/. Contains the two ways of writing the same result, 6 answers out of 270 and 4 combinations out of 90, together with the question-by-question detail.
- HENUSSE Barometer: https://henusse.com/en/barometer/. Contains the reading of 11 August 2026 and the published 0 out of 30 for our own name.
- Mary L. McHugh, “Interrater reliability: the kappa statistic”, Biochemia Medica, 2012: https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/. Contains Jacob Cohen’s 1960 criticism of a simple agreement rate, which cannot account for agreement obtained by chance, and the proposal of a corrected coefficient.
- EQUATOR Network, STARD 2015 checklist for reporting diagnostic accuracy studies, Bossuyt et al., BMJ 2015;351:h5527: https://www.equator-network.org/reporting-guidelines/stard/. Establishes the existence of a published list of required items for reporting on an instrument that produces a verdict. The checklist itself, available as a PDF, contains 30 items, including item 1 on stating the measure of accuracy, items 10a and 10b on the test being evaluated and its standard, items 6 to 9 on the population tested, item 16 on missing results and item 26 on limitations.
All the addresses cited above were opened and checked on 16 August 2026.
Page published on 16 August 2026. The studies and readings cited each carry their own date; engines change quickly, and a figure on this page says nothing about a later measurement. If one of the cited documents changes, this page will be corrected and the correction dated.