How do you measure whether an AI visibility fix worked?
AI answers vary run to run, engine to engine, region to region. Without a baseline, a stable prompt set and repeated scans, you cannot tell a win from noise. Here is how to measure honestly.
This is the verification stage of the getting cited by AI pillar, and it is what separates an optimization program from a guessing program.
What should be measured before making a change?
The baseline: for each buyer question you care about, which engines answered, whether your brand was mentioned, whether it was recommended, which URLs were cited, and which competitors appeared. Collected more than once, because a single run cannot show you the variance you will later need to distinguish from improvement.
Baselines are cheap insurance. Every "did it work?" argument three months from now is settled by whether you recorded what the answers looked like before. A change shipped without a baseline can never be credited with anything.
What is the difference between mentioned, recommended and cited?
Three separate outcomes that move independently, and conflating them corrupts every measurement:
- Mentioned: your brand name appears anywhere in the answer, in any role, including "unlike X" or a neutral listing.
- Recommended: the answer presents you as a choice for the buyer's need, which is the outcome revenue actually cares about.
- Cited: one of your URLs appears as a supporting source, which can happen without a mention and vice versa.
A fix can move one without the others: a rewritten page can earn citations while the recommendation still goes to a competitor with stronger reviews. Tracking the three separately tells you which stage of the journey moved; a single blended score hides exactly the information you need.
Why should results be tracked prompt by prompt?
Because fixes are prompt-specific and averages bury them. A page fix targets particular questions; if you measure a portfolio-wide visibility number, a real win on three fixed questions drowns in the noise of two hundred untouched ones, and an unrelated loss elsewhere can mask it entirely.
Prompt-level tracking also produces the diagnosis for the next round: which questions you win, which you lose, and who wins the ones you lose. That per-question evidence is what makes the next recommendation grounded instead of generic.
How many engines and locations should be included?
Every engine your buyers plausibly use, and every region you actually serve. Coverage decisions are measurement decisions: engines select different sources for the same question, and answers differ by region, so a fix can work on one engine in one country and be invisible in your dashboard if you only test another.
For local businesses, region is not optional: "best electrician near me" answers are location-dependent by construction, and measuring from the wrong region measures a different question. Test from where your buyers are.
Why is one AI answer not a reliable trend?
Because generation is not deterministic. The same question to the same engine on the same day can produce different answers with different sources; across days, retrieval changes as pages update and indexes refresh. One answer is one sample from a distribution.
The failure mode this produces is real: someone checks ChatGPT the morning after a fix, sees the brand mentioned, and declares victory; a week later the mention is gone and the fix is declared a failure. Both conclusions came from single samples. Repeated runs are the only way to see whether the distribution moved.
How should a before-and-after test be designed?
Five properties make the test survivable:
- A stable prompt set. The same questions, phrased the same way, before and after. Changing the questions mid-test destroys comparability.
- A recorded baseline across multiple runs, so the pre-change variance is known.
- One change at a time per affected question set, with the ship date noted. Three simultaneous changes produce an unattributable result.
- The same engines and regions in every run.
- A decision window agreed in advance: how many post-change scans you will look at before judging, so the judgment is not made on whichever single run looks best.
Recrawl and answer refresh vary by engine, site and page; there is no universal number of days after which the result is in. State the mechanism honestly: the change improves your likelihood of being selected, and repeated retesting verifies whether it did.
What should happen when the answer does not change?
Diagnose which stage failed, in order, rather than shipping more of the same fix. Confirm the changed page is actually being served and indexed (crawler access first). Check whether the engines' cited sources for the question even include pages like yours, or whether the answer is grounded in reviews and threads your page cannot influence; that finding redirects effort to independent proof.
Sometimes the honest conclusion is that the question is dominated by sources you are not on yet, and the on-page work was necessary but not sufficient. That is the citation journey working as designed: eligibility first, corroboration second, outcome probabilistic throughout. The diagnostic article walks the failure classes in order.
How CiteAgentic links recommendations to later scans
Product example. Every CiteAgentic recommendation is tied to the specific buyer questions it should affect. The scan history is the baseline; when a fix is approved and shipped, the affected questions are re-tested on subsequent scans, and the page state, validation state and AI outcome are reported separately: the page now answers the question, independent corroboration exists or is still missing, and the answer outcome moved from not present to mentioned, cited or recommended, or did not move yet. No part of the product claims a fix will produce a citation; it records whether one did.
Baseline your questions this week
Start tracking your real buyer questions across engines and regions now, so every future fix has a before to compare against. Start free trial →