Skip to content

How many AI answers before the result means anything.

A business named in 3 of 10 AI answers could have a true naming rate from 11% to 60% (95% confidence, fresh sessions, same wording, one engine). Real answers vary more than that, so the true range is wider. By the formula below, at most 97 answers to one question bring the range within plus or minus 10 points.

Check your count

Use this if you asked an AI assistant one question a few times and want to know what 2 of 5 proves. SparkToro found the same whole list of brands in fewer than 1 in 100 pairs of ChatGPT or Google AI answers (November and December 2025).

What do you want to check?
Examples

How it decides

If you doubt a result, a spreadsheet can recompute it, except plan mode's exact chances and headline count (our own calculation).

The rules this calculator applies, in order, with the source of each
StepRuleWhere it comes from
InputsAnswers collected: a whole number from 1 to 1,000. Answers that named you: a whole number from 0 up to answers collected. Rates for the planning check: 0% to 100%, and the two must differ. Anything else gets a message, not a result.This calculator.
Rate and 95% rangep = named ÷ n and z = 1.96. Low = (p + z²/(2n) − z × √(p × (1 − p)/n + z²/(4n²))) ÷ (1 + z²/n). High is the same with + before z. If named is 0, low is 0%. If named equals n, high is 100%. All shown in whole percent.Wilson score interval. Wikipedia prints the formula and says it can be used with small samples. Brown, Cai and DasGupta (2001) recommend it for small n.
Verdict on one rateWidth = high − low, in the whole points shown. More than 40: too few answers to say. 21 to 40: a rough range. 20 or less: a usable range, about plus or minus 10 points.Our own cut-offs.
Answers needed for a marginWrite m and r as fractions: 10 points is m = 0.10, and 50% is r = 0.5. For a margin m at a true rate r, n = z² × r × (1 − r) ÷ m², rounded up. Worst case, r = 50%: margins of 5, 10, 15, 20 points need 385, 97, 43, 25 answers. At r = 10% and m = 10 points, 35 answers.The standard formula for estimating a proportion, as on Wikipedia (its width W is 2m).
Did it change?p1 before and p2 after, each with its Wilson bounds: l1 and u1 for p1, l2 and u2 for p2. Change d = p2 − p1. Low = d − √((p2 − l2)² + (u1 − p1)²). High = d + √((u2 − p2)² + (p1 − l1)²). If the whole-point range excludes 0, the change is clear. If it includes 0, it cannot be told from run-to-run variation.Newcombe (1998) reports that combining the two Wilson intervals performs well. The abstract does not print the formula. Our code is checked against the statsmodels library's newcomb method (z = 1.96): the same whole-point results on 3,000 random pairs of counts.
Textbook answers to see a changeWrite the rates as fractions (20% is 0.20) and p̄ = (p1 + p2) ÷ 2. For a change from p1 to p2, n per period = (1.96 × √(2 × p̄ × (1 − p̄)) + 0.8416 × √(p1 × (1 − p1) + p2 × (1 − p2)))² ÷ (p1 − p2)², rounded up. 1.96 sets a 5% false-alarm rate and 0.8416 an 80% chance of a clear result. If n × the smaller of p̄ and 1 − p̄ is under 5, the textbook count is flagged as unreliable.The customary normal-approximation formula, as on the University of British Columbia calculator, which also warns when expected counts are under 5.
Headline answers to see a changeThe card headlines the first count per period from which the chance of a clear result stays at 80% or more, up to 1,000 answers. Whole counts make the chance uneven: for 0% to 10% it falls from 87% at 74 answers to 77% at 75. The first count to touch 80% is therefore not used. If no count up to 1,000 gets there, the card says so and shows the chance at 1,000.Our own rule.
Exact chances at a countThe card adds up, over every pair of counts, the binomial chance of the pair when the check calls it clear. For the clear chance the true rates are p1 then p2, and only a clear change in the true direction counts. For the false-alarm chance the true rate stays at p̄ in both periods, and a clear rise or a clear fall counts.Our own calculation, not checkable in a spreadsheet. A second program written separately gave the same figures.
Worked examples to recompute in a spreadsheet
CheckCount enteredResult
One rate0 of 40% to 49%: too few answers to say
One rate2 of 512% to 77%: too few answers to say
One rate3 of 1011% to 60%: too few answers to say
One rate6 of 2015% to 52%: a rough range
One rate0 of 200% to 16%: a usable range, and zero stays inside it
One rate30 of 10022% to 40%: a usable range
Change4 of 20, then 8 of 20+20 points, range −8 to +44: no clear change
Change4 of 50, then 20 of 50+32 points, range +16 to +47: a clear rise
Textbook count20% to 40%82 per period, 164 in total
Textbook count20% to 30%294 per period, 588 in total
Exact chances of the Did it change? check. Our own calculation: it cannot be redone in a spreadsheet
True ratesAnswers per periodChance of a clear resultFalse-alarm chance at the average rate
20% to 40%82 (textbook formula)79%4%
20% to 40%85 (headline count)80%4%
20% to 30%294 (textbook formula)76%4%
20% to 30%325 (headline count)80%3%
1% to 3%769 (textbook formula)56%1%
1% to 3%1,000 (the most this calculator accepts)68%1%
0% to 10%74 (textbook formula)87%2%
0% to 10%7577%2%
0% to 10%78 (headline count)80%2%
224 pairs of rates from 5% to 95% in 5-point steps, gaps of 10 points or more, no unreliable flagThe textbook count for each pair75.5% to 85.3%3.2% to 6.4%

What it cannot tell you

  • Independence. Logged-in repeats can share history, so the true range is wider.
  • Buyer wording. SparkToro's 142 prompts for one task barely resembled each other, yet returned a consistent set of brands.
  • Cause. A name is not a recommendation, and a count does not say why. For visitors, see how to see whether AI sends you visitors and what AI assistants cite for local businesses.
  • Many comparisons. Reporting only the question that moved makes false alarms likely.

Sources

Sources for every external figure and formula on this page, with date and who publishes each
SourceWhat it saysDateWho publishes it
SparkToro, AIs are highly inconsistent when recommending brands600 volunteers ran 12 prompts 2,961 times on ChatGPT, Claude and Google AI. For ChatGPT and Google's AI, fewer than 1 in 100 pairs of answers gave the same list of brands, and fewer than 1 in 1,000 the same order. The author advises asking at least 60 to 100 times to learn an assistant's set of recommendations. Of 142 prompts that volunteers wrote for one task, barely two looked similar to the author. Yet those prompts returned a consistent set of brands: top headphone brands appeared in 55% to 77% of 994 answers. The author calls visibility percentage across dozens to hundreds of prompts, each run several times, a reasonable metric. How many runs are enough is listed as an open question.Published January 27, 2026. The runs took place in November and December 2025, and the post says newer models may differ. Read October 6, 2026.SparkToro sells audience-research software. The study was run with a co-author who works at a company that sells AI visibility tracking; the post says so.
BrightLocal, local AI visibility studyIn repeated runs of the same local prompt, a business that appeared once was seen again in about half of later attempts. The rates were 50% on ChatGPT and Google AI Mode and 58% on Google AI Overviews.Published September 16, 2026. Read October 6, 2026.BrightLocal sells local SEO software, including an AI visibility tracker.
Wikipedia, Binomial proportion confidence intervalThe Wilson score interval formula, and that it can be used with small samples.Page last edited September 11, 2026. Read October 6, 2026.Wikipedia, a nonprofit encyclopedia. Sells nothing.
Brown, Cai and DasGupta, Interval estimation for a binomial proportionThe authors recommend the Wilson interval, or the Jeffreys interval, for small n.Statistical Science, May 2001. Read October 6, 2026.A statistics journal. Sells no AI tracking.
Newcombe, Interval estimation for the difference between independent proportionsAbstract: a method combining Wilson score intervals for the two proportions “performs well, and is readily implemented irrespective of sample size”. The abstract does not print the formula.Statistics in Medicine 17(8): 873 to 890, April 30, 1998. Read October 6, 2026.A medical statistics journal. Sells no AI tracking.
statsmodels, confint_proportions_2indepLists newcomb as a method for the interval of the difference between two proportions, and cites Newcombe (1998) among its references. It does not print the formula.No date shown. Read October 6, 2026.The statsmodels developers, an open-source Python library. Sells nothing.
Wikipedia, Sample size determinationThe sample size for estimating a proportion, with 0.5 as the most conservative rate.Page last edited July 22, 2026. Read October 6, 2026.Wikipedia, a nonprofit encyclopedia. Sells nothing.
Rollin Brant, two-proportion sample size calculatorThe sample size per group for two proportions, with 5% two-sided significance and 80% power by default, and a warning when expected counts are under 5.No date shown. Read October 6, 2026.A statistics page on the University of British Columbia's statistics department site. Sells nothing.
Montrelia AI visibility scoreboardMontrelia was named in 0 of 4 buyer questions in each of two runs on Ask Brave. The scoreboard reports counts, never percentages.Runs on September 28 and October 2, 2026.Montrelia, which sells GEO. One engine, our own measurement.

Questions people ask

How many times should I run a ChatGPT prompt to check brand visibility?

SparkToro's author advises asking at least 60 to 100 times to learn an assistant's set of recommendations (January 27, 2026). By this page's formula, plus or minus 10 points needs up to 97 answers at a naming rate near 50%, and about 35 near 10%. Use a fresh session and the same wording each time.

Is AI visibility tracking accurate?

It depends on the answers behind each percentage, so ask any tool how many it rests on and whether each was a fresh session. SparkToro's author found 142 differently worded prompts still returned a consistent set of brands, and calls a brand's share of appearances across many runs a reasonable metric. The same author rejects ranking position in AI answers.

What does being named in 0 of 5 AI answers mean?

It means the true rate could still be as high as 43% (95% range, fresh sessions, one engine). Zero of 100 would still allow up to 4%. A zero never proves a business is never named; more answers only lower the ceiling.

How do I check whether a website change moved my AI mentions?

Collect answers with the same questions, wording, engine and place before and after, then use the Did it change? check. A rise from 20% to 40% takes about 85 answers in each period, 170 in total. This page cannot say how long an engine takes to pick up a site change, so it cannot say when to start the second count.

Last updated October 6, 2026

Talk to us for 15 minutes.