Measuring visibility
Why one answer misleads
AI assistants sample their answers. Ask ChatGPT the same question twice in a row and the list of recommended products can change order, gain a brand or drop one. A single run that mentions your brand says it can appear; it says little about how often it does. Brand visibility is a rate: out of many answers to the same prompt, the share that mention you, and where they place you. In our study of 596 answers, two thirds of the brands ChatGPT named for a question appeared in some runs and not others.
How many samples
A rate from a handful of runs is a wide range. The table shows the 95% Wilson interval for a brand mentioned in 60% of its runs: the range the true rate plausibly lies in, given only that many answers.
| Mentioned | Measured rate | 95% interval |
|---|---|---|
| 3 of 5 | 60% | 23% to 88% |
| 6 of 10 | 60% | 31% to 83% |
| 12 of 20 | 60% | 39% to 78% |
| 30 of 50 | 60% | 46% to 72% |
| 60 of 100 | 60% | 50% to 69% |
Take at least 10 samples per prompt and engine, and spread them over hours: answers drawn in one burst share whatever the engine was doing that minute. When comparing two brands or two weeks, read a difference into the rates only when their intervals are clearly apart.
With plain HTTP
Submit the prompt N times as one batch (up to 500 items), each item with its own idempotencyKey. Deriving the key from the prompt, market,
day and sample number makes a resubmission after a timeout safe: a key already used is refused instead of running twice.
curl -X POST "https://api.answerline.dev/v1/async/task/batch" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '[{"taskType":"CHATGPT","payload":{"prompt":"Which CRM is best for a 10-person sales team?","country":"US"},"idempotencyKey":"crm-chatgpt-us-2026-09-28-01"},{"taskType":"CHATGPT","payload":{"prompt":"Which CRM is best for a 10-person sales team?","country":"US"},"idempotencyKey":"crm-chatgpt-us-2026-09-28-02"},{"taskType":"CHATGPT","payload":{"prompt":"Which CRM is best for a 10-person sales team?","country":"US"},"idempotencyKey":"crm-chatgpt-us-2026-09-28-03"}]'
Poll each task until it is COMPLETED or FAILED, or receive it by webhook:
curl -X GET "https://api.answerline.dev/v1/async/task/b27a21e1-7c39-4aa2-a347-23e828c426f9" \ -H "Authorization: Bearer $API_KEY"
{
"task": {
"id": "b27a21e1-7c39-4aa2-a347-23e828c426f9",
"taskType": "CHATGPT",
"status": "COMPLETED",
"priority": 1,
"createdAt": "2026-09-28T09:00:00.000Z",
"latencyMs": 18250,
"idempotencyKey": "crm-chatgpt-us-2026-09-28-01"
},
"credits": {
"creditsToCharge": 5,
"creditsCharged": 5
},
"response": {
"success": true,
"result": {
"text": "For a 10-person team, Globex is the usual pick for pipeline reporting. Acme CRM is simpler to set up and costs less per seat, which matters more at that size.",
"sources": [
{
"position": 1,
"url": "https://www.crm-reviews.example/best-small-team",
"label": "Best CRMs for small teams",
"description": "…"
}
]
}
}
}
Then count: a run mentions the brand when response.result.text contains one of its names as a whole word, ignoring case. The mention rate is
mentioned runs over completed runs; failed tasks are left out of both, and are not charged. Cited sites are in sources[].url and, on engines
that return them, citationPills[].url.
With the SDKs
visibility.measure does the above: it submits samples tasks per engine (default 10, at most 100), each under a fresh key,
optionally in rounds spaced apart, waits for them and returns per engine requested, completed, failed,
the results and, given a brand, their brand check. brandCheck / brand_check runs the same check on results
you already have, one engine at a time.
import { Client } from "@answerline/sdk";
const client = new Client({ apiKey: process.env.API_KEY! });
const measured = await client.visibility.measure({
prompt: "Which CRM is best for a 10-person sales team?",
engines: ["chatgpt", "gemini"],
country: "US",
samples: 12,
rounds: 4,
spacingMs: 2 * 60 * 60 * 1000, // 4 rounds of 3, two hours apart
brand: { name: "Acme CRM", aliases: ["Acme"], domains: ["acme.example"] },
competitors: [{ name: "Globex" }, { name: "Initech" }],
});
for (const [engine, m] of Object.entries(measured)) {
const { mentionRate, interval, averageRank } = m.brandCheck!.brand;
console.log(engine, m.completed, "of", m.requested, mentionRate, interval, averageRank);
}
import os
from answerline import Client
with Client(os.environ["API_KEY"]) as client:
measured = client.visibility.measure(
"Which CRM is best for a 10-person sales team?",
["chatgpt", "gemini"],
"US",
samples=12,
rounds=4,
spacing_seconds=2 * 60 * 60, # 4 rounds of 3, two hours apart
brand={"name": "Acme CRM", "aliases": ["Acme"], "domains": ["acme.example"]},
competitors=[{"name": "Globex"}, {"name": "Initech"}],
)
for engine, m in measured.items():
b = m["brand_check"]["brand"]
print(engine, m["completed"], "of", m["requested"], b["mention_rate"], b["interval"], b["average_rank"])
{
"name": "Acme CRM",
"runs": 12,
"mentioned": 7,
"mentionRate": 0.583,
"interval": {
"low": 0.32,
"high": 0.807
},
"averageRank": 1.714,
"ownCitations": 3,
"mentions": [
{
"run": 0,
"rank": 2,
"matches": [
{
"text": "Acme CRM",
"snippet": "For a 10-person team, Globex is the usual pick for pipeline reporting. Acme CRM is simpler to set up and costs less per seat, which matters more at that size.",
"field": "text"
}
]
},
"…"
],
"citedWhenAbsent": [
{
"domain": "crm-reviews.example",
"runs": 4
},
{
"domain": "salesops.example",
"runs": 2
},
{
"domain": "globex.example",
"runs": 1
}
]
}
Python returns the same fields in snake_case (mention_rate, average_rank, own_citations,
cited_when_absent). Each sample is one async task, priced like any other; see credits and billing.
Reading the numbers
| Field | Meaning |
|---|---|
mentionRate, interval | Share of answers naming the brand or an alias, as a whole word in the text or in entities, with its 95% Wilson interval. |
averageRank | Where the brand first appears among the brands you checked, 1 being first, averaged over the answers that mention it. |
ownCitations | Answers citing at least one URL on the brand's domains or their subdomains. |
mentions[].matches | Every match with up to 80 characters of context either side, to check each count by eye. |
citedWhenAbsent | The 10 sites cited most often in answers that leave the brand out: the pages shaping those answers. |
Limits
Samples taken minutes or hours apart measure how much an engine varies its own answer to one prompt from one country. They do not reproduce what different people see: real users phrase questions differently, carry conversation history and personalization, and use other markets. Treat a mention rate as a tracked indicator for a fixed prompt set, compared across weeks and engines. Matching is by name: a brand whose name is a common word is also counted where the word is used in its ordinary sense, so read the snippets, and add aliases for product names and spellings the answers use.