HENUSSE Making your company readable by AI. Outside, and in.

Understand · Searching or answering from memory

When does ChatGPT actually search the web, and when does it answer from memory?

An answer engine does not consult the web for every question. Question by question, it decides whether to search for pages or answer with what it retained during training. We counted. On the night of 14 to 15 August 2026, across the thirty frozen questions in a client’s public diagnostic, the ChatGPT application triggered its search for only eleven of them. Ten of those eleven were buying questions.

This page explains what happens at that moment, and what a business leader can do with it. Every claim refers to a provider’s documentation, published research or a dated reading whose conditions are public.

For which questions does the machine search?

The rule fits into one sentence: it searches when the question concerns something that may have changed, and answers from memory when it judges the question to be stable.

We measured this split instead of assuming it. On the night of 14 to 15 August 2026, in the ChatGPT application, we asked one by one the thirty buyer questions frozen in advance for a client’s public diagnostic. The application triggered its web search for only eleven of them: the ten buying questions, in which a buyer compares offers, plus one question about a language examination. For the nineteen general advisory questions, it answered without citing a single source, and nobody appeared as a source, neither our client nor its competitors. The details, screenshots and limitations are published in our correlation study.

The scope remains attached to the figure: one website, thirty questions, one night, one account. We draw no general law about ChatGPT from it, and nobody should draw one from a reading of this size. We draw one observation, which we will repeat in the next reading.

This observation has a direct commercial consequence. In this reading, the questions for which the machine searched were precisely those that decide a purchase: comparing offers, choosing a provider, deciding between two solutions. For those questions, the answer is composed from the pages found, and a company absent from those pages cannot appear as a source. In other words, your presence in the sources governs your presence in the answer. The reverse does not follow, however: appearing among the pages found gives no right to be used. Making a company readable and verifiable by these machines is our work, and the rules we impose on ourselves when measuring it are written in our measurement charter.

The providers’ documentation describes the same mechanism in the tools they provide to developers. OpenAI writes that the model can choose whether or not to search the web according to the content of the question asked. Its answer then includes by default the links to the pages found. Anthropic writes that Claude decides whether to search based on the question. It searches when the request depends on current, changing information or information outside its training data, and answers directly when the request concerns stable knowledge. These documents are written for developers and do not describe the inside of consumer applications, which no provider publishes. That is one more reason to measure the application itself as well.

What does the machine have before it when it answers you?

Everything the machine has before it fits into a single space for text, known in the field as the context window. Anthropic defines it as all the text the model can consult to produce its answer, including that answer. The provider specifies that it is working memory, separate from the large corpus on which the model was trained.

It contains your conversation and what the model’s tools place there: the pages returned by a search, the document you attached, the result of a calculation. All of this is read like the rest of the conversation. Anything that is not there does not exist for the machine at the moment it writes.

Why is everything read again with each message?

A model retains no memory of its own from one request to the next. OpenAI writes that each generation request is independent and stateless, and that a multi-turn conversation is possible only because the previous messages are supplied to it again. Providers do offer systems that avoid sending everything again. OpenAI documents server-side storage of the exchange. Anthropic notes that chat interfaces such as claude.ai can manage the window on a rolling basis by removing the oldest messages. These systems change the plumbing, not the principle. What the machine processes to write its next answer is the text accumulated since the start of the exchange.

Two consequences follow, and a business leader encounters both. The first is mechanical: the longer the conversation, the more text there is to process at each turn, so every answer takes more time and resources. With the same content, the same message uses more computation at the end of a long discussion than in a new thread, and the wait reflects this.

The second is more troublesome: attention becomes diluted. Anthropic writes that the model’s accuracy and ability to recall information decline as the accumulated text grows longer, and calls this phenomenon context rot, literally the degradation of the context. Its engineers describe an attention budget that every new addition of text consumes.

This finding is not merely a provider’s opinion. In 2023, Nelson F. Liu and his colleagues published “Lost in the Middle: How Language Models Use Long Contexts”. Their measurement establishes that performance is highest when the useful information appears at the beginning or the end of the supplied text. It declines markedly when the model must find that information in the middle of a long context. This work concerns the models available in 2023, and today it serves as a warning: the middle of a long conversation is the place where an instruction is most likely to be lost.

The working memory of a conversation A diagram in two parts. At the top, four columns show that at each turn the whole previous exchange is read again: the stack of text has one band at the first turn, two at the second, three at the third and four at the fourth; at the fourth turn, the oldest band, at the bottom of the stack, is shown with a dotted outline because it is eventually summarised or removed. At the bottom, a schematic curve shows attention: high at the beginning of the supplied text, low in the middle, then high again at the end. The dip in the middle is marked in red. That is where an instruction is most likely to be lost. FIGURE 5 · THE WORKING MEMORY OF A CONVERSATION At each turn, the whole previous exchange is read again turn 1 turn 2 turn 3 turn 4 the oldest, summarised or removed What you wrote earlier still affects the answer you receive today, until the machine removes it, without telling you. Where attention declines in a long text ATTENTION beginning middle end of the text supplied to the machine The middle of a long exchange is where an instruction is most likely to be lost. Schematic curve. The phenomenon was measured in 2023 by Liu and his colleagues in “Lost in the Middle”, using the models of that period. Anthropic calls this decline context rot and describes an attention budget.

Why is the tool that measures not necessarily what your customer sees?

An answer in an application depends on the account asking the question. A measurement instrument must ask the question in a new session, with no history, so that a third party can repeat the reading and reproduce the same counts. The two surfaces can therefore diverge, and we found no public documentation from any provider describing the connection between the tool supplied to developers and the consumer application.

We measured this gap instead of debating it. During the same night of 14 to 15 August 2026, the thirty questions were asked on both sides. The two surfaces returned the same presence verdict thirty times out of thirty when the majority verdict from the instrument’s three passes was compared with the application’s one pass. They agree twenty-nine times out of thirty under the strictest rule, in which a single pass is enough to count a presence. The comparison rule, screenshots and limitations are published in the study itself, which this page does not repeat.

This finding does not settle the question, and we do not present it as doing so. In August 2026, the Paris company H Company published a comparison between OpenAI’s application programming interface and the ChatGPT web application, using buying queries. It concludes that the two behave like different search engines. It finds that about 18% of the products highlighted by the application also appear in the application programming interface’s answers, and that the overlap between the cited sources is about 2.5%. Both findings hold at the same time because they concern different objects, as our study explains. A business leader should retain this: a visibility figure means nothing without the surface tested and the date of the reading. That is the rule we apply to our own figures, including when they are zero, as shown by our barometer.

What should you do with all of this in your company?

Three habits are enough to avoid the most common mistakes.

  • Start a new conversation for a new subject. Everything that came before still affects the answer until the machine removes it. A short thread containing the right documents is better than an endless thread containing everything.
  • Provide the documents instead of hoping that the machine knows them. What you paste enters its working memory. What you assume it knows may not be there at all, especially when it concerns your figures, prices or conditions.
  • Be wary of an answer that arrives at the end of a long discussion. An instruction given at the beginning ends up where attention declines most, as the work cited above shows. Ask the question again in a new thread, then compare the two answers.

Frequently asked questions

How can I tell whether the machine searched or answered from memory?

Look at the links. When a search has taken place, OpenAI’s and Anthropic’s documentation says that the answer displays the sources found by default. An answer with no link at all was therefore, in all likelihood, composed from what the model retained during training. In our reading, the nineteen answers with no search cited no source.

Does ChatGPT read my website live when someone asks a question?

Only if it triggers a search, and only if your pages are among the results returned by that search. When it does not search, it answers with what it retained during training, and no page is consulted or cited. The useful question therefore becomes one of access: do the engines’ robots have the right and the ability to read your website?

Does this split of eleven questions out of thirty apply to my sector?

There is no basis for saying that it does. This split was observed on one website, thirty questions and one night in August 2026. The only honest way to know what happens with your own questions is to freeze them, ask them and count, on a written date.

Why does a long conversation eventually forget my instructions from the beginning?

Because its attention is distributed across an increasingly long text, and because the start of an exchange is eventually removed or summarised. The work published in 2023 and cited above shows that information placed in the middle of a long context is used least effectively. Repeating the instruction when it needs to apply costs less than hoping that it has held.

Can a model that searches still be wrong?

Yes, because searching is not verifying. Search reduces the risk by requiring the answer to rely on real pages and displaying the links, which allows you to check the source yourself. It does not turn the machine into an authority, and a displayed link remains a link that needs to be opened.

Sources

We accessed and checked these providers’ documents, this research and this publication on 16 August 2026.

  • Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang, “Lost in the Middle: How Language Models Use Long Contexts”, arXiv:2307.03172, submitted on 6 July 2023, last revised on 20 November 2023: https://arxiv.org/abs/2307.03172. This work establishes that performance is highest when the useful information appears at the beginning or end of the context, and that it declines markedly in the middle of a long context. This measurement was conducted using the language models available in 2023.
  • Anthropic, “Context windows”: https://platform.claude.com/docs/en/build-with-claude/context-windows. This page defines the context window as working memory separate from the training corpus, including the answer. It describes the accumulation of turns, names the decline associated with length, and notes that chat interfaces such as claude.ai can manage the window on a rolling basis.
  • Anthropic, “Effective context engineering for AI agents”, 29 September 2025: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents. This text describes context as a finite resource and speaks of an attention budget that each addition of text consumes.
  • Anthropic, “Web search tool”: https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool. This page says that Claude decides whether to search based on the question, that it searches when the request depends on current, changing information or information outside its training data, that it answers directly when the request concerns stable knowledge, and that citations are always enabled for web search.
  • OpenAI, “Web search” guide: https://developers.openai.com/api/docs/guides/tools-web-search. This guide says that the model can choose whether or not to search the web according to the content of the question. It specifies that the answer includes by default citations to the addresses found by the search.
  • OpenAI, “Conversation state” guide: https://developers.openai.com/api/docs/guides/conversation-state. This guide says that each generation request is independent and stateless, and that previous messages must be supplied again to maintain a conversation. It also describes server-side storage of the exchange, which avoids sending them again.
  • H Company, “Learning From the User”, August 2026: https://hcompany.ai/geo-using-computer-use-agents-vs-api-limitations. This study compares OpenAI’s application programming interface with the ChatGPT web application using buying queries. It finds that about 18% of the products highlighted by the application reappear in the application programming interface’s answers, that the overlap between citations is about 2.5%, and concludes that the two surfaces behave like different search engines.

Here are our own pages, cited above.

The house figures cited here come from a single reading, dated the night of 14 to 15 August 2026, covering one website and thirty questions. The next one is planned for November 2026, and the differences from this one will be published.

This page was published on 16 August 2026. The cited documents may change; their access date appears above. If one of them changes its position, this page will be corrected and the correction dated.

WhatsApp