Check your count
Use this if you asked an AI assistant one question a few times and want to know what 2 of 5 proves. SparkToro found the same whole list of brands in fewer than 1 in 100 pairs of ChatGPT or Google AI answers (November and December 2025).
How it decides
If you doubt a result, a spreadsheet can recompute it, except plan mode's exact chances and headline count (our own calculation).
| Step | Rule | Where it comes from |
|---|---|---|
| Inputs | Answers collected: a whole number from 1 to 1,000. Answers that named you: a whole number from 0 up to answers collected. Rates for the planning check: 0% to 100%, and the two must differ. Anything else gets a message, not a result. | This calculator. |
| Rate and 95% range | p = named ÷ n and z = 1.96. Low = (p + z²/(2n) − z × √(p × (1 − p)/n + z²/(4n²))) ÷ (1 + z²/n). High is the same with + before z. If named is 0, low is 0%. If named equals n, high is 100%. All shown in whole percent. | Wilson score interval. Wikipedia prints the formula and says it can be used with small samples. Brown, Cai and DasGupta (2001) recommend it for small n. |
| Verdict on one rate | Width = high − low, in the whole points shown. More than 40: too few answers to say. 21 to 40: a rough range. 20 or less: a usable range, about plus or minus 10 points. | Our own cut-offs. |
| Answers needed for a margin | Write m and r as fractions: 10 points is m = 0.10, and 50% is r = 0.5. For a margin m at a true rate r, n = z² × r × (1 − r) ÷ m², rounded up. Worst case, r = 50%: margins of 5, 10, 15, 20 points need 385, 97, 43, 25 answers. At r = 10% and m = 10 points, 35 answers. | The standard formula for estimating a proportion, as on Wikipedia (its width W is 2m). |
| Did it change? | p1 before and p2 after, each with its Wilson bounds: l1 and u1 for p1, l2 and u2 for p2. Change d = p2 − p1. Low = d − √((p2 − l2)² + (u1 − p1)²). High = d + √((u2 − p2)² + (p1 − l1)²). If the whole-point range excludes 0, the change is clear. If it includes 0, it cannot be told from run-to-run variation. | Newcombe (1998) reports that combining the two Wilson intervals performs well. The abstract does not print the formula. Our code is checked against the statsmodels library's newcomb method (z = 1.96): the same whole-point results on 3,000 random pairs of counts. |
| Textbook answers to see a change | Write the rates as fractions (20% is 0.20) and p̄ = (p1 + p2) ÷ 2. For a change from p1 to p2, n per period = (1.96 × √(2 × p̄ × (1 − p̄)) + 0.8416 × √(p1 × (1 − p1) + p2 × (1 − p2)))² ÷ (p1 − p2)², rounded up. 1.96 sets a 5% false-alarm rate and 0.8416 an 80% chance of a clear result. If n × the smaller of p̄ and 1 − p̄ is under 5, the textbook count is flagged as unreliable. | The customary normal-approximation formula, as on the University of British Columbia calculator, which also warns when expected counts are under 5. |
| Headline answers to see a change | The card headlines the first count per period from which the chance of a clear result stays at 80% or more, up to 1,000 answers. Whole counts make the chance uneven: for 0% to 10% it falls from 87% at 74 answers to 77% at 75. The first count to touch 80% is therefore not used. If no count up to 1,000 gets there, the card says so and shows the chance at 1,000. | Our own rule. |
| Exact chances at a count | The card adds up, over every pair of counts, the binomial chance of the pair when the check calls it clear. For the clear chance the true rates are p1 then p2, and only a clear change in the true direction counts. For the false-alarm chance the true rate stays at p̄ in both periods, and a clear rise or a clear fall counts. | Our own calculation, not checkable in a spreadsheet. A second program written separately gave the same figures. |
| Check | Count entered | Result |
|---|---|---|
| One rate | 0 of 4 | 0% to 49%: too few answers to say |
| One rate | 2 of 5 | 12% to 77%: too few answers to say |
| One rate | 3 of 10 | 11% to 60%: too few answers to say |
| One rate | 6 of 20 | 15% to 52%: a rough range |
| One rate | 0 of 20 | 0% to 16%: a usable range, and zero stays inside it |
| One rate | 30 of 100 | 22% to 40%: a usable range |
| Change | 4 of 20, then 8 of 20 | +20 points, range −8 to +44: no clear change |
| Change | 4 of 50, then 20 of 50 | +32 points, range +16 to +47: a clear rise |
| Textbook count | 20% to 40% | 82 per period, 164 in total |
| Textbook count | 20% to 30% | 294 per period, 588 in total |
| True rates | Answers per period | Chance of a clear result | False-alarm chance at the average rate |
|---|---|---|---|
| 20% to 40% | 82 (textbook formula) | 79% | 4% |
| 20% to 40% | 85 (headline count) | 80% | 4% |
| 20% to 30% | 294 (textbook formula) | 76% | 4% |
| 20% to 30% | 325 (headline count) | 80% | 3% |
| 1% to 3% | 769 (textbook formula) | 56% | 1% |
| 1% to 3% | 1,000 (the most this calculator accepts) | 68% | 1% |
| 0% to 10% | 74 (textbook formula) | 87% | 2% |
| 0% to 10% | 75 | 77% | 2% |
| 0% to 10% | 78 (headline count) | 80% | 2% |
| 224 pairs of rates from 5% to 95% in 5-point steps, gaps of 10 points or more, no unreliable flag | The textbook count for each pair | 75.5% to 85.3% | 3.2% to 6.4% |
What it cannot tell you
- Independence. Logged-in repeats can share history, so the true range is wider.
- Buyer wording. SparkToro's 142 prompts for one task barely resembled each other, yet returned a consistent set of brands.
- Cause. A name is not a recommendation, and a count does not say why. For visitors, see how to see whether AI sends you visitors and what AI assistants cite for local businesses.
- Many comparisons. Reporting only the question that moved makes false alarms likely.
Sources
| Source | What it says | Date | Who publishes it |
|---|---|---|---|
| SparkToro, AIs are highly inconsistent when recommending brands | 600 volunteers ran 12 prompts 2,961 times on ChatGPT, Claude and Google AI. For ChatGPT and Google's AI, fewer than 1 in 100 pairs of answers gave the same list of brands, and fewer than 1 in 1,000 the same order. The author advises asking at least 60 to 100 times to learn an assistant's set of recommendations. Of 142 prompts that volunteers wrote for one task, barely two looked similar to the author. Yet those prompts returned a consistent set of brands: top headphone brands appeared in 55% to 77% of 994 answers. The author calls visibility percentage across dozens to hundreds of prompts, each run several times, a reasonable metric. How many runs are enough is listed as an open question. | Published January 27, 2026. The runs took place in November and December 2025, and the post says newer models may differ. Read October 6, 2026. | SparkToro sells audience-research software. The study was run with a co-author who works at a company that sells AI visibility tracking; the post says so. |
| BrightLocal, local AI visibility study | In repeated runs of the same local prompt, a business that appeared once was seen again in about half of later attempts. The rates were 50% on ChatGPT and Google AI Mode and 58% on Google AI Overviews. | Published September 16, 2026. Read October 6, 2026. | BrightLocal sells local SEO software, including an AI visibility tracker. |
| Wikipedia, Binomial proportion confidence interval | The Wilson score interval formula, and that it can be used with small samples. | Page last edited September 11, 2026. Read October 6, 2026. | Wikipedia, a nonprofit encyclopedia. Sells nothing. |
| Brown, Cai and DasGupta, Interval estimation for a binomial proportion | The authors recommend the Wilson interval, or the Jeffreys interval, for small n. | Statistical Science, May 2001. Read October 6, 2026. | A statistics journal. Sells no AI tracking. |
| Newcombe, Interval estimation for the difference between independent proportions | Abstract: a method combining Wilson score intervals for the two proportions “performs well, and is readily implemented irrespective of sample size”. The abstract does not print the formula. | Statistics in Medicine 17(8): 873 to 890, April 30, 1998. Read October 6, 2026. | A medical statistics journal. Sells no AI tracking. |
| statsmodels, confint_proportions_2indep | Lists newcomb as a method for the interval of the difference between two proportions, and cites Newcombe (1998) among its references. It does not print the formula. | No date shown. Read October 6, 2026. | The statsmodels developers, an open-source Python library. Sells nothing. |
| Wikipedia, Sample size determination | The sample size for estimating a proportion, with 0.5 as the most conservative rate. | Page last edited July 22, 2026. Read October 6, 2026. | Wikipedia, a nonprofit encyclopedia. Sells nothing. |
| Rollin Brant, two-proportion sample size calculator | The sample size per group for two proportions, with 5% two-sided significance and 80% power by default, and a warning when expected counts are under 5. | No date shown. Read October 6, 2026. | A statistics page on the University of British Columbia's statistics department site. Sells nothing. |
| Montrelia AI visibility scoreboard | Montrelia was named in 0 of 4 buyer questions in each of two runs on Ask Brave. The scoreboard reports counts, never percentages. | Runs on September 28 and October 2, 2026. | Montrelia, which sells GEO. One engine, our own measurement. |
Questions people ask
How many times should I run a ChatGPT prompt to check brand visibility?
SparkToro's author advises asking at least 60 to 100 times to learn an assistant's set of recommendations (January 27, 2026). By this page's formula, plus or minus 10 points needs up to 97 answers at a naming rate near 50%, and about 35 near 10%. Use a fresh session and the same wording each time.
Is AI visibility tracking accurate?
It depends on the answers behind each percentage, so ask any tool how many it rests on and whether each was a fresh session. SparkToro's author found 142 differently worded prompts still returned a consistent set of brands, and calls a brand's share of appearances across many runs a reasonable metric. The same author rejects ranking position in AI answers.
What does being named in 0 of 5 AI answers mean?
It means the true rate could still be as high as 43% (95% range, fresh sessions, one engine). Zero of 100 would still allow up to 4%. A zero never proves a business is never named; more answers only lower the ceiling.
How do I check whether a website change moved my AI mentions?
Collect answers with the same questions, wording, engine and place before and after, then use the Did it change? check. A rise from 20% to 40% takes about 85 answers in each period, 170 in total. This page cannot say how long an engine takes to pick up a site change, so it cannot say when to start the second count.
Last updated October 6, 2026
