Skip to main content

How we measure the Ozvor AI Visibility Score

The Ozvor AI Visibility Score is one umbrella number. It’s made up of three distinct sub-scores: Visibility, Citation Readiness, and Execution. Each measures something different, and each is something you can act on. Here is exactly how each is computed.

Honest methodology — measured signals only, baselines labelled as such. Every number tells you where it came from: either a live probe of your brand, or a neutral placeholder.

Three sub-scores, one umbrella

The three sub-scores are not replacements for each other — they answer three different questions about your brand’s position in AI search.

Visibility

What AI engines see

Based on real AI probes — how often, how high, how well.

Citation Readiness

How ready you are to be cited

Site signals + entity authority you can directly improve.

Execution

How much you've done

Live progress on your prioritised action plan.

1. Visibility Score

What AI engines actually do with your brand when a buyer asks about your category. This is the only score that requires real AI API calls and cannot be gamed by changing your website.

Visibility

What AI engines see when asked about your brand

Sourced from live probes across 5 AI engines. AI answers vary by day and engine.

Citation Rate50% of score

Fraction of buyer questions where an AI engine mentioned your brand. We ask each buyer question in 2 different wordings, and we run each wording 2 times per engine. That is 4 runs per question, per engine, before we add anything. From those runs we compute a mention rate, not a single coin flip.

Average Position Score30% of score

When cited, how high in the AI answer? Position 1 = 1.0, position 2 = 0.5, position 3 = 0.33, and so on. Zero if never cited.

Sentiment Score20% of score

We classify the text around each brand mention as positive, neutral, or negative using a deterministic phrase-matching classifier. Positive = 1.0, neutral = 0.5, negative = 0.

Visibility = clamp((citationRate×0.50 + avgPositionScore×0.30 + sentimentScore×0.20)×100, 0, 100)

Why we run each question more than once

AI engines are not consistent. The same question can get a different answer on the next request. So we never judge your brand on one run.

We start lean. Each buyer question is asked in 2 wordings, and each wording runs 2 times per engine. That is a base of 4 runs per question, per engine.

Then we only spend more where the answer is unclear. If a question lands in the grey zone on an engine, meaning your brand is cited in 25% to 75% of the runs, we add 1 more run per wording. We do that for at most 2 extra rounds, and we stop once that question reaches 6 runs on that engine. A clean 0 of 4 or 4 of 4 stays at 4 runs. Paying for more runs would not tell you anything new.

Every audit has a ceiling. There is a hard limit on how many AI runs one audit can spend (220 by default). If an extra round would cross it, we stop adding runs instead of cutting corners elsewhere. The base 4 runs always finish.

Every rate ships with its margin of error. We report a 95% Wilson confidence interval next to each rate, so you can see how solid it is. 4 of 4 runs gives a tight range. 2 of 4 gives a wide one, and that width is the honest part. When only 1 run exists (older audits), we show “single sample” instead of claiming confidence we do not have.

Honest note: AI outputs are not deterministic. The same prompt can produce different answers on different days. We run each buyer question at least 4 times per engine, add runs where the signal is unclear, and publish a confidence interval with every rate. It is still a snapshot. Re-audit weekly to track the trend, rather than treating any single score as absolute.

2. Citation Readiness Score

How ready your brand is to be cited by AI engines. Unlike the Visibility Score, these are signals you can directly control and improve — your site, your entity presence, your off-site authority.

Citation Readiness

How ready your content and presence is to be cited

Derived from two internal vectors: Performance (60%) and Brand (40%).

CitationReadiness = round(Performance×0.60 + Brand×0.40)

Performance vector (60% of Citation Readiness)

Technical visibility and citation share — things AI crawlers look at before they decide whether to cite you.

Citation share-of-voice vs competitors30% of vector

Your citation rate relative to the brands you benchmarked against.

Schema.org coverage30% of vector

Are your pages marked up with standard structured data? Standard schema markup helps AI engines understand your pages. No special AI-only schema is required, per current SEO guidance.

AI-crawler access25% of vector

What fraction of AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) can access your site?

Multi-page content citation-worthiness15% of vector

Do your pages use statistics, sourced claims, quotations, and question-answer structure that AI engines are trained to cite?

Brand vector (40% of Citation Readiness)

Entity authority and off-site presence — how the broader web and knowledge graphs represent your brand.

Entity completeness (Wikidata/Wikipedia)40% of vector

How complete is your brand’s entry in Wikidata/Wikipedia? We check official website, industry, founding date, description, LinkedIn, and Crunchbase links.

Citation volume40% of vector

Raw mention count across the AI engines we probe, normalised to 0–1.

E-E-A-T signal (Reddit, G2, LinkedIn, off-site authority)20% of vector

Blends your Reddit footprint (threads, subreddits, sentiment), off-site authority (presence on the 7 highest-AI-cited sources: Reddit, Wikipedia, LinkedIn, G2, Trustpilot, Crunchbase, YouTube), and on-site identity signals from a live crawl of your homepage.

Honest note: these are the signals you can directly control and improve. Schema, crawler access, entity records, and off-site presence are all editable. The Execution score tracks how many of the recommended fixes you have completed.

3. Execution Progress

After your audit, Ozvor generates a prioritised action plan of tasks (plan_task cards). Execution tracks how far through that plan you are. It is the only score that is fully in your hands.

Execution Progress

% of your recommended action cards completed

After the audit is complete, Ozvor generates a prioritised list of action cards — one per gap identified. Execution is the percentage of non-rejected cards you have marked done.

Execution = (done cards / total non-rejected cards) × 100
No cards yetNot started

If no action cards exist yet, we show “Not started” — not a 0% bar. We never fabricate a number. Run an audit and generate your plan to start tracking.

Cards in progress0–100%

As you mark cards done, the percentage updates live. Rejected cards are excluded from the denominator so declining a low-priority fix doesn’t punish your score.

Honest note: this is a live counter, not an audit snapshot. It updates immediately when you complete or reject action cards. A high Execution score means you have done the work — and a re-audit will confirm whether Visibility improved as a result.

The 5 AI engines we probe

We query each engine through its official API. We never scrape web interfaces or browser sessions.

EngineAPI used
ChatGPT (GPT-4o)OpenAI Chat Completions API
Claude (claude-3-5-sonnet)Anthropic Messages API
PerplexityPerplexity API (returns citation URLs natively)
GeminiGoogle Generative Language API
Google AI OverviewDataForSEO SERP API (captures real SERP results including AI Overviews)

Our methodology commitment

We query each AI engine through its official API and record whether your brand is cited, where in the answer it appears, and how it is described. We never scrape LLM web interfaces. We repeat each probe multiple times and report a mention rate — not a single test — so results are more reliable.

Every score shows “Measured” (real signal from this audit) or “Baseline” (a neutral placeholder shown transparently when a data source is not yet connected) so you always know exactly how confident each number is.

The Execution Progress score is the one number we will never estimate or interpolate — it is always the exact ratio of completed cards to total non-rejected cards, and is shown as “Not started” (not 0%) until your first action plan exists.

What changed in version 2.1

Our methodology carries a version number, and every audit records the version that produced it. Version 2.1 changes what counts as a citation. Here is exactly what changed and why your number may move.

Methodology 2.1

A mention is not the same thing as a citation

Until version 2.0, we counted a citation any time your brand name appeared in an AI answer. That was too generous. A name in a sentence is not always a recommendation.

Now we read every answer twice

The first pass finds every place your brand is named and quotes the exact words. The second pass is a blind check. A separate reviewer reads the raw answer and one candidate mention at a time, with no idea what the first pass decided, then rules on whether that mention is really about you and what kind of mention it is.

Only two kinds count toward your score

Direct recommendationCounts

The answer actually suggests your brand to the person asking.

Cited sourceCounts

The answer leans on your site or your content as the source of what it says.

Neutral mentionDoes not count
Negative mentionDoes not count

Neutral and negative mentions still show up in your report. You can read every one of them. They simply do not earn points.

Four things the blind check catches

  1. Same name, different company. An answer about “Acme Corp” the spring factory is not about Acme the software company. It no longer counts for you.
  2. The answer says no. “I would not recommend Acme for this” used to score the same as praise. It is now a negative mention, and it does not count.
  3. Your brand only inside a link. If your name appears only in a URL and nothing in the text recommends you, it is logged as a cited source, not as a recommendation. It still counts, but it is now labelled for what it is, so a link is never dressed up as an endorsement.
  4. A name in a list, with no opinion attached. Being listed in a comparison table with no endorsement is a neutral mention, and it does not count.

Your number may drop. That is the point.

If your score falls after this change, your brand did not get worse. We stopped counting things that were never citations. The old number was inflated by false positives. The new number is smaller and truer, and it is the one you can actually build on.

The safety rule behind it: the second pass can only remove a citation, never invent one. Nothing is added to your score by this change.

Scores from 2.0 and 2.1 are not comparable. Treat your first 2.1 audit as a new baseline, and compare 2.1 against 2.1 from there.

Every audit shows its version. Each audit stores the methodology version that produced it, and the report shows it. When we tighten the rules again, the version goes up and you will see it. You never have to guess which rules made your score.

Methodology version history

VersionDateWhat changed
1.0LaunchFlat repeat protocol. Each buyer prompt was asked a fixed number of times per engine, and any appearance of the brand name counted as a citation.
2.028 July 2026Intent based prompt portfolio with sequential sampling: a lean base of runs per wording, extra runs added only where the result was ambiguous, and a Wilson 95% confidence interval reported on every aggregate. All 5 engines probed on their search enabled surfaces.
2.129 July 2026Two pass extraction with a blind verifier. Only verified mentions of the kinds direct recommendation and cited source count as a citation. Neutral and negative mentions are reported but do not score.

What we measure vs. what’s coming

Measured nowStill on the roadmap
ChatGPT, Claude, Perplexity, Gemini citationsReal-time daily monitoring (requires API budget)
Google AI Overview (via SERP API)Bing AI, other emerging engines
Reddit presence (threads, subreddits, sentiment)Quora, niche forums
Wikidata/Wikipedia entity consistencyCrunchbase, LinkedIn automated cross-check
On-site schema, crawler accessFreshness / content age signals
Multi-page content citation-worthinessAuto-schema generation
Action plan execution tracking (live counter)Automated verification of completed fixes
How We Measure the Ozvor AI Visibility Score | Ozvor | Ozvor