Notes

How long does it take for a website change to show up in AI answers?

You changed your website and you want to know when AI assistants will reflect that change. The frustrating answer is that a website change can appear in a web-grounded answer within days while the model’s stored understanding takes months to change.

Two clocks are running

People often ask for one waiting period because traditional search encouraged that expectation. A page was crawled, indexed and eventually displayed for a query, so a change in visibility could be discussed as one moving process.

AI assistants can assemble an answer through several routes. Some answers use a current web search, some rely on information retained during training and some combine both without showing you which parts came from where.

That observation creates two different clocks. The retrieval clock concerns whether a system can find your updated page today. The model-memory clock concerns whether a future response treats the updated description as part of its general knowledge.

Two routes to an answer, two different waiting periods
RouteWhat has to changeTypical lagCan you observe it
Answer that searches the webThe page is fetched and indexed againDaysYes, by testing with search enabled
Answer from stored knowledgeThe model itself is retrained or refreshedMonthsNo, not directly
The practical consequence is that one number cannot answer the question. An assistant that searched the web may reflect your change this week while the same assistant, answering from memory, repeats the old description for months. Most published advice quietly addresses only the top row.

Those clocks may agree eventually, but they do not move together. Your updated page can sit there perfectly available to a crawler while an answer produced without web retrieval carries on using the old description.

This is a structural difference rather than something we have measured over time and it is worth being precise about which is which. Retrieval indexes refresh continuously, so a changed page can reach a web-grounded answer as soon as it is fetched again. Stored knowledge changes only when the model itself is retrained or refreshed and for frontier models those events are separated by months rather than days. Nobody outside the provider can watch that happen, which is exactly why a single promised timeframe should be treated with suspicion.

What a website change can influence quickly

The first useful question is narrow. Can the new information actually be fetched, read and attributed to the right organisation by something that is not a person or does it exist only for a human with a browser?

Plenty of pages fail well before any model gets as far as making a judgement. In one measured run, a site served 321KB of HTML while only 4.2 percent was readable text. That does not prove that the site was ignored, but it demonstrates how a technically successful page load can still provide very little usable description.

What a crawler could read on one measured page Of 321 kilobytes of HTML delivered, about 13.5 kilobytes or 4.2 percent, was text a crawler could read. The remaining 95.8 percent was markup, scripts and styling. HTML delivered 321 KB Text a crawler could read 13.5 KB, 4.2%
Measured on one real page, which returned 321KB of HTML with about 4.2 percent of it readable text. Everything else was markup, script and styling that carries no description of the organisation. Throughout, the site looked entirely normal to a human visitor, which is what makes this failure easy to miss.

Check whether the important statement appears in ordinary page text. Confirm that the page is accessible without an interaction that a crawler cannot complete. Make the organisation’s name, category, audience and location clear enough that another source can identify the same entity.

These changes may affect answers that search the web. They cannot guarantee that an assistant will select your organisation, even when the page is retrieved.

Retrieval does not guarantee selection

Finding a page is only one stage of producing an answer. The assistant must decide whether the page supports the question, whether the organisation fits the implied need and whether another source appears more relevant.

One recorded run produced 41 citations across the tested questions. Every citation came from four questions that named the organisation directly, while questions describing a need produced none.

The result matters because it separates two common diagnoses. A company may be absent because its information cannot be retrieved. It may also be retrievable only when the user supplies the name, which means the system has not connected the organisation to the broader need.

Test both kinds of question during each review. Use a named prompt that asks about the organisation directly. Then use a neutral prompt that describes the problem, audience or buying situation without naming any provider.

Record whether the answer cites the relevant page, mentions the organisation and explains why it belongs in the answer. A citation on a named question shows access. It does not show recognition in an open recommendation.

Memory changes are harder to schedule

You cannot set the retraining calendar for a frontier model. You also cannot inspect its internal memory and confirm that a particular sentence has been replaced.

That makes many claims about permanent improvement untestable in the short term. If an assistant answers from stored knowledge, a revised page may remain absent until a later training or refresh event. The timing may differ by model, topic and source prominence.

A new page might therefore produce an immediate web-grounded change without producing a stable general answer. Conversely, a model might repeat an old description after the page has changed because that description remains present elsewhere.

The honest position is not that waiting is pointless. It is that the waiting period has an upper uncertainty that no website owner can remove.

What a timing test has to separate

The design problem is not complicated, but it is easy to get wrong in a way that produces a confident and meaningless answer.

A comparison needs a recorded starting point, because without one you are comparing today against a memory of what the answers used to look like. That memory is unreliable and it tends to flatter whichever change you just made.

The two conditions then have to be kept apart. An answer produced with search available and an answer produced without it are measuring different clocks and pooling them gives a number that moves for reasons you cannot attribute. If a result shifts only when search is available, that is evidence about retrieval and nothing has yet been learned about memory.

The last requirement is repetition, because identical asks can produce different results. A single favourable answer is not a milestone and treating it as one is how a team convinces itself that a change worked.

What the test cannot tell you

This process can show that an answer changed after a page change. It cannot prove that the page caused every change, because ranking, retrieval, prompt interpretation and model updates may also have shifted.

It can estimate how often the new description appears under the conditions you tested. What stays invisible is every other conversation, whatever retrieval systems sit behind the interface and any future training data.

The next useful action is to split your monitoring into retrieval and memory questions. Track whether the updated page can be found now, then track whether the assistant uses the updated description without being given the organisation’s name.

What remains uncertain is the date when a general model will absorb the change, if it does so at all. That uncertainty is a property of the system, not evidence that your website work has failed.

AnswerFit measures how AI models describe, compare and recommend one organisation, then re-tests after a change to see whether the answer moved. The methodology page sets out how that measurement is put together and what it deliberately refuses to claim.