A response is an observation
Traditional search reports encouraged precise-looking measurements. A page could be assigned a position for a query, then compared with its previous position under similar conditions.
An assistant answer is less like a fixed position and more like one observation from a changing process. The wording may shift, the sources may differ and the model may choose a different organisation even when the prompt is identical.
In one measured run, the same question produced non-recognition and then accurate recognition. Neither answer should be treated as the complete truth about the organisation’s visibility.
| Ask | What the model said | Conclusion if this were your only check |
|---|---|---|
| First | Did not recognise the organisation | Absent from model knowledge and work is needed |
| Second | Described it accurately | Present and correctly understood, so nothing to do |
The immediate consequence deserves careful measurement. A score based on one run has no visible uncertainty, even though uncertainty is present. You need repeated observations before you can discuss a pattern.
Decide what visibility means first
The word visibility hides several different outcomes. You might care about being named, being cited, being described correctly or being recommended for a specific customer need.
Those outcomes should not be collapsed into one number at the start. An organisation that is named in an answer may still be described incorrectly. Another may be described accurately but omitted from a recommendation because the question does not fit its offer.
Write the target outcome in observable terms. For example, you could measure the percentage of need-based prompts where the organisation is named and the reason given is factually correct.
Add separate fields for citation, category accuracy, competitor relevance and whether the answer asks the user to verify the claim. Clear definitions make later disagreement useful instead of confusing.
Build questions around real needs
Your question set determines what your measurement can see. If every prompt includes the organisation’s name, the test mostly measures whether the assistant can retrieve information after being handed the answer’s subject.
That is a legitimate diagnostic, but it does not measure discovery. Add questions that describe the problem, audience, location and constraints without naming a provider.
Include questions with different levels of specificity. A broad need reveals category association. A detailed buying situation tests fit. A comparison prompt tests whether the assistant places the organisation beside relevant alternatives.
Avoid writing every question in the same voice. Real people use short prompts, incomplete context and different descriptions of the same need. Your test should represent the language that could reasonably lead to a recommendation.
One warning deserves particular emphasis here. A question set that scores everything perfectly may be measuring its own design. One auto-drafted set graded all 38 of its questions at 4 out of 4, with the result that two different headline measures reported an identical figure.
That result does not establish excellent visibility. It establishes that the test lacked enough difficulty or independence to distinguish outcomes.
Repeat the same prompts carefully
A fixed prompt set makes change observable. Vary the wording between runs and the comparison quietly stops meaning anything, because a changed answer becomes indistinguishable from a changed question.
Whatever conditions surround the ask are part of the measurement rather than incidental detail, so they have to travel with the result. An answer recorded without knowing which model produced it or whether search was available, cannot be compared against anything later.
Keeping the raw answers matters more than it sounds. Compression throws away the sentence that would have told you the assistant reached the right conclusion for the wrong reason and that sentence is often the useful part.
You can report a proportion with a range rather than a single exact claim. A small sample will still be noisy, but its uncertainty is visible. A larger sample improves the estimate without making the underlying system fully predictable.
Keep grading separate from ranking
Human judgement enters as soon as you decide whether an answer is accurate or relevant. Some of that is unavoidable and the right response is to make the choice explicit rather than to pretend it was not made.
Use a short rubric with defined bands. One band might mean no recognition, another partial recognition, another accurate recognition without recommendation and the highest band accurate recognition with a relevant recommendation.
Ask more than one reviewer to grade a subset when the stakes justify it. Record disagreements instead of smoothing them away. A disagreement may show that the rubric needs a clearer boundary.
The grader can also change the headline number. Two model families were asked to grade the same 38 questions and one sat at or below the other on every single question, with a mean gap of 0.79 bands. Their spread was almost identical, which points to a difference in severity rather than a disagreement about which questions were stronger.
This distinction is useful for decisions. Treat ranking as a signal about relative weakness, while treating the exact band as an estimate affected by the grading method.
Compare changes with a baseline
Before changing any website copy, freeze a baseline and keep it. Save the prompts, the raw answers, the grading rules and the conditions the test ran under. Then change one meaningful group of inputs at a time, such as the organisation description or a single page explaining a service, so that a later difference has one candidate cause rather than several.
Wait long enough for web retrieval to encounter the change, while marking any known model refresh or interface change. Run the same prompts again and compare both outcome frequency and answer content.
A before-and-after difference is informative, but it is not a perfect causal experiment. Other websites may have changed, retrieval indexes may have shifted and the assistant may have changed internally.
If you can, include prompts that should not be affected by the edit. They provide a modest control for broad changes in the testing environment. The control will not isolate every cause, yet it is stronger than comparing two unrelated prompt sets.
What your number cannot claim
Your measurement can describe observed behaviour under recorded conditions. It cannot represent every user, every prompt or every private assistant session.
Such a result still cannot prove that a website edit permanently changed model knowledge. An untracked conversation also cannot become a known referral or sales opportunity through this method.
The next action is to define one visibility outcome, create a balanced prompt set and run repeated baseline observations. Report the rate, the sample size, the conditions and the main disagreements together.
What remains uncertain is how well those observations generalise beyond the tested assistants and questions. A careful measurement system does not remove that uncertainty. It gives you enough structure to stop confusing one answer with a trend.