Skip to main content
Main content
EdenRank Blog

Experiments and evidence

The Control Group Playbook for AI Citation Content Tests

Choose a control group for AI citation testing: freeze the estimand, match on pre-test citation rate, and label results with an evidence tier.

EdenRank Editorial TeamPublished Aug 12, 20269 min read
The Control Group Playbook for AI Citation Content Tests: A precise editorial timeline of routed evidence displaying the physical divergence between cited source cards and uncited.

In brief

  • For this protocol, freeze the citation metric - numerator, denominator, and exclusions - and the control group selection before collecting any post-intervention data.
  • Match control pages on pre-existing citation rate and content type; run the same fixed query set on treatment and control under identical session conditions.
  • Report results using a prespecified evidence class (Level 1, 2, or 3), explicitly listing the confounders the design cannot exclude.
Sections in this article

TL;DR

  • The failure: before-and-after charts without control groups cannot attribute citation changes to content decisions.
  • Play 1: Freeze the metric - define citation rate with a fixed denominator, numerator, and exclusions before testing.
  • Play 2: Match controls - select pages with similar pre-test citation visibility and content type, not random ones.
  • Play 3: Pre-register - write and date a document that locks in your design, timeline, and analysis plan.
  • Play 4: Collect data - run queries with a scripted method and log citations in a structured, auditable format.
  • Play 5: Label evidence - assign a Level 1, 2, or 3 class and state what your design can and cannot prove.
9 min read

Key takeaways

Freeze your estimand, denominator, and exclusion rules in a prespecification document before touching any content.

Match control pages on pre-test citation rate and content type - not domain authority or random selection.

Set the observation window length based on your topic's query volatility and known crawl delays, not a universal rule.

Use difference-in-differences to subtract background citation movement from your treatment effect.

Label every result with an evidence class (Level 1, 2, or 3) and state what the design cannot prove.

A single controlled observation is one data point; accumulate multiple tests before drawing directional conclusions.

Why Your Before-and-After Citation Chart Fails Audit

o choose a control group for an AI citation content test, freeze your estimand first - the exact metric you want to attribute to the content change - then match control pages on pre-existing citation visibility before you touch any content. Without that sequence, your before-and-after chart is noise dressed as evidence. The rest of this playbook is the protocol.

The core failure mode is familiar: an operator updates a set of product pages with better sourcing, runs the same queries some weeks later, sees more citations, and calls it a win. A simple time series of your own pages cannot separate those forces. You need a matched control group that experienced the same external environment while receiving no treatment.

Denominator table: what changes the count at each level

Unit of observationNumeratorDenominatorCommon exclusions
Content template (e.g., all comparison pages)Queries where any URL in the template group is citedTotal queries × number of URLs in the groupQueries that resolve to a different template type, duplicate-URL appearances in one answer
DomainQueries where any domain URL is citedTotal queries in the fixed query setQueries where the domain appears only in a non-citation context (e.g., inline mention without a source link)

Play 1: Freeze the Citation Metric and Its Denominator Before You Choose a Control

Before you select a single control page, write down the unit of observation, the numerator definition, the denominator definition, and every exclusion rule. This is your prespecification document. Its purpose is to prevent you from adjusting any of these definitions after you see the data.

Start with the unit: are you measuring a single URL, a content template, or a domain? Each choice produces a different denominator. Next, build a fixed query set of non-branded informational queries that your content could plausibly answer. The size of that set is a design decision you make based on the topic coverage of your test pages - there is no universal number. Exclude brand-navigation queries where you already dominate and queries that return no AI-generated answer at all. Each query in the set is one opportunity, so your denominator is the count of queries multiplied by the number of pages if you test multiple pages at the URL level.

Define what counts as a citation: an exact URL match in the AI answer's source list, or a brand mention with a link? For content tests, exact URL match is the cleaner signal because brand mentions vary by provider and session state. Document the session conditions - signed-out, no personalization, consistent locale - and record them in the prespecification document.

A/B Testing 101 - NN/G states: "A/B testing can help UX teams determine the improvements in the user experience that are best for their business goals ."

Operator checkpoint

Your metric is citation rate: citations to a target URL divided by opportunities (queries) tracked. Lock the formula before you run a single query.

Play 2: Match Control Units on Pre-Existing Citation Visibility, Not Domain Authority

You need matched pairs: for each page you intend to change (treatment), select a control page that had a similar citation rate in the baseline period, covers a comparable subtopic, and was published in a similar timeframe.

Run your fixed query set against both groups before any content change and record the citation count for each page. Calculate the average citation rate and its distribution for each group. If the control group's average rate is materially higher than the treatment group's, your post-test comparison will be biased upward - you will understate the treatment effect. If it is materially lower, you will overstate it. Rebalance the groups until the pre-treatment rates are close. If you cannot find adequate matches within your own site, pause the test and expand your topic coverage first rather than proceeding with a mismatched control.

Record the matching criteria in your prespecification document: which pages are in each group, why they were selected, and what the pre-treatment citation rates were. This record is what makes the design auditable.

Matching trap

Matching on domain authority alone ignores the page-level citation rate. A high-DA domain with zero pre-test citations is not a valid control for a page that was already cited regularly.

See whether EdenRank fits the workflow

Request access to the Citation OS and review the process, evidence boundaries, and plan fit. Browse all free tools

Request access

Play 3: Pre-Register the Treatment, Assignment, and Observation Window

Pre-registration is the single step that separates a credible controlled observation from a post-hoc rationalization. Write every parameter of your test in a document, date it, and share it with at least one other person before you make any content change. The document must include: the treatment description (exactly what you will change and on which pages), the control group list, the fixed query set, the providers you will query, the primary metric formula, the exclusion rules, and the planned analysis method.

The timeline structure matters. After the baseline period, assign control and treatment units formally (record the date). Make the content change. Then wait a post-treatment observation period before running the post-intervention queries.

You set this window based on your knowledge of the providers you are testing; there is no universal duration that applies across all providers and topics.

After the observation window closes, run the analysis exactly as prespecified. Do not add queries, swap pages, or change the citation definition. If you discover a data-collection error, document it and report it as a limitation rather than silently correcting it.

  1. Write the prespecification document with all parameters listed above and record the date
  2. Run the fixed query set on both groups and record the baseline citation count per page
  3. Verify that treatment and control groups have comparable pre-treatment citation rates; rebalance if not
  4. Run the fixed query set on both groups again using the same session conditions as baseline
  5. Calculate the difference-in-differences: (treatment post − treatment pre) − (control post − control pre)

Checklist

  • Confirm that make the content change on treatment pages only; record the exact date of each change
  • Confirm that wait the pre-specified post-treatment observation window before collecting post-intervention data
  • Confirm that report the result with the evidence class label and the limitations stated in the prespecification document

Operator checkpoint

Pre-registration prevents retroactive adjustment of the hypothesis, control group composition, or success metric after seeing results. A dated document is the cheapest insurance you have.

Play 4: Run Queries, Record Citations, and Calculate the Rate

Executing the measurement requires a scripted approach to avoid manual inconsistencies. Use the exact same prompt format and session state for every query in every collection period.

A screenshot folder is not a structured log - it cannot be queried or audited. If you automate collection, record the script version and any API parameters alongside the data so the collection method is reproducible.

Calculate citation rate for each group and period: citations divided by opportunities. Then calculate the difference-in-differences. A synthetic illustration (labeled as such): if treatment pages gained eight citations and control pages gained three citations over the same period, the adjusted treatment effect is five - not eight. The control group corrected for three citations that appeared regardless of your content change. Real data will not be this clean; document the actual numbers and their variance.

Log field minimum

Every row in your citation log needs: query text, provider, model version, collection date, target URL, citation present (yes/no), session state, and anomaly flag. Missing any field makes the record unauditable.

Play 5: Label the Result with an Evidence Class, Not a Causal Claim

Assign an evidence class before you share the result. The class communicates what the design can support and what it cannot. Three levels cover most content test designs: Level 1 is an uncontrolled observation - a before-and-after on a single group with no control. It carries no causal weight and should be labeled as such. Level 2 is a controlled observation - a matched control group with a pre-post design and a difference-in-differences calculation. It suggests attribution but cannot rule out selection bias or unmeasured confounders.

Level 3 is a randomized experiment - random assignment of content to treatment and control with a holdout, which is rare in content testing but produces the strongest attributable evidence.

Prespecify the evidence class you are aiming for before the test. That is a defensible result for a skeptical stakeholder. What you cannot do is run a Level 1 design and report it as if it were Level 2.

Your pre-registration document should state the evidence class, the known limitations, and the confounders the design cannot exclude. When you share the result, lead with the evidence class label before the number.

Level 1 vs Level 2

Before

Level 1 - Uncontrolled observation: single group, before-and-after only. Cannot separate content effect from external changes. Report as: 'Citation count changed; cause unknown.'

After

Level 2 - Controlled observation: matched control group, pre-post design, difference-in-differences. Suggests attribution; cannot rule out selection bias. Report as: 'Adjusted treatment effect was X citations in a controlled observation with stated limitations.'

References and further reading

These links are provided for direct inspection. A reference is not treated as proof of every statement in this article.

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.

Written by

EdenRank Editorial Team

The product and editorial team documents repeatable ways to inspect AI-answer visibility, source evidence, and content operations.

8References
ShownMethod
1Evidence claims

Expertise

AI answer visibility measurementCitation & source intelligenceLLM readiness & crawlabilityEntity trust & schema markupPrompt strategy & buyer signals

Published

Aug 12, 2026

About EdenRankAll articles

Want insights like this for your own brand?

Talk to the team

Published by EdenRank.