Three tests get bundled together
People who ask this usually have one of three jobs in mind. Some want to know how assistants behave in a specific country. Others want to know what happens when the question is asked in another language. A third group wants to know what a model they host themselves currently says about the organisation.
Those jobs share a method and still need separate samples, because the answers are not repeats of the same test. You write questions a real person could ask, send them under recorded conditions, keep the raw answers and grade whether the organisation was named, described accurately and recommended for a need it actually serves.
A public assistant answering in English from a default location is one surface. The same intent asked from another country, in another language or against a private endpoint is a different surface. Combining those answers into one visibility score treats unlike observations as if they were repeats.
The practical limit follows from that design. A measurement tool can only evaluate a model it can call, under conditions it can record. Coverage that was never collected is a gap in the sample rather than a missing trick of the method.
Regional measurement needs its own sample
Regional behaviour appears for two reasons that should stay distinct. The assistant may retrieve different web results when the request appears to come from another country. The market may also be served by a different product altogether, popular there and unavailable through the interfaces a typical English-language panel already covers.
Country therefore has to travel with the result. Record the intended market, the location settings used for the request, whether web retrieval was on and the date. Keep that group’s denominator separate from every other country.
The common failure is to add a country label after the fact and then report a blended rate. An organisation named often in one market and omitted in another will look moderately visible in the average, which is exactly the number that stops anyone acting on either market.
Location itself is an imperfect signal, because a proxy, a language header or an account setting can influence retrieval without proving that a person in that country would see the same answer. Record how location was asserted, and treat unconfirmed location as a reason to exclude the run from the regional sample rather than as a slightly noisier data point.
Some markets are served by assistants that never appear in a default panel. Those have to be added as their own models. Inferring them from a location setting on a different product will measure the wrong assistant.
Start with the markets that actually affect pipeline. Adding every country at once multiplies cost and variation before you know whether the first market’s pattern is stable enough to compare against.
Another language is not a translated score
Sending the same intent in another language is technically straightforward. Grading the result as if it were the English test is where the measurement usually goes wrong.
A useful multilingual set is written in the buying language of that market. Literal translation preserves English sentence structure and often the English category names, which real buyers may never type. The retrieval layer then searches a different web, so the sources an assistant can cite are not a translated copy of the English sources either.
The answer language can also diverge from the request language. An assistant might reply in English to a non-English prompt, mix both or drift out of the requested language. Those outcomes belong in the record, because a rate that ignores them will treat a language failure as ordinary variation.
Grading is the other weak point. A reviewer or judge that is fluent in one language and weak in another will shift the headline number even when the underlying answers are similar. Severity differences of that kind already appear between grader families in a single language, and they appear between languages too.
Keep per-language scores with their own denominators. A translated average should not be promoted as the organisation’s multilingual visibility, because it erases the language in which the organisation is actually failing.
A locally hosted model answers a different question
A model running on your own machines can often be asked through the same kind of interface used for public APIs. The measurement method still applies: fixed questions, repeated asks, preserved answers and an explicit rubric.
What changes is the claim you are allowed to make. A local result tells you what that private assistant currently says, which cannot be read as what buyers hear when they use the public consumer assistants that already sit in the market.
Web retrieval is the usual missing piece. Many local deployments answer from stored weights only, which is a test of what the model appears to know already. That result is useful for an internal assistant and a weak proxy for web-grounded discovery, because the public answer a buyer sees is often assembled from live search rather than from weights alone.
If the local setup does add search, the index is still not the index used by a public answer engine. Citations from the two setups cannot be compared as if they came from the same web. The comparison remains informative about your internal tool while remaining a separate test from the public surface.
Privacy is often the reason the model is local in the first place. Sending those prompts through a hosted measurement service can undo that choice. The work can stay on the same network as the model, provided the answers, grades and run conditions are still kept in a form you can inspect later.
Include a local model when staff, partners or customers actually use it. Leave it out of the public-visibility sample when they do not, so a private result cannot inflate or deflate the number you use for buyers.
What to keep in separate cells
A cell is the smallest unit that still means something: one country, one query language, one model or hosting arrangement and one retrieval mode. Each cell keeps its own questions, its own denominator, its own baseline and its own exclusions.
| Surface | What a result can support | What mixing it with other cells hides |
|---|---|---|
| A public assistant asked from a stated country | Whether that product, under those location settings, named or recommended the organisation | A different regional assistant, or a blended global rate that looks moderate while one market is empty |
| The same intent written in another language | Whether the organisation is recognised in the language buyers actually use | The English-language rate, and the fact that the retrieved sources are a different web |
| A model hosted on your own network | What that private assistant currently says, from its weights and any local search you attached | What public buyers hear, and how a public web index cites you |
If a run comes back in the wrong language, from an unconfirmed location or from a different model than the one requested, exclude it from that cell’s rate and record why. Inclusion would make the cell look more complete than the evidence is.
What remains worth doing
Pick the surfaces your buyers actually use. Write a small question set for one market and one language first, then repeat the asks instead of stretching immediately into a matrix you cannot grade well.
When you add a second language or a local endpoint, freeze the first cell as a baseline rather than folding the new answers into it. Compare within a cell over time. Treat a difference between cells as a difference between tests, because the conditions that produced the answers were never the same.
What remains uncertain is how closely any asserted location or translated prompt matches a real person in that market, and how long a local model’s stored knowledge will lag the public web. A careful setup leaves that uncertainty visible, which is the condition for not reading the wrong surface as coverage of the right one.