Experiments and evidence
The Control Group Playbook for AI Citation Content Tests
Choose a control group for AI citation testing: freeze the estimand, match on pre-test citation rate, and label results with an evidence tier.

In brief
- For this protocol, freeze the citation metric - numerator, denominator, and exclusions - and the control group selection before collecting any post-intervention data.
- Match control pages on pre-existing citation rate and content type; run the same fixed query set on treatment and control under identical session conditions.
- Report results using a prespecified evidence class (Level 1, 2, or 3), explicitly listing the confounders the design cannot exclude.
Sections in this article
TL;DR
- The failure: before-and-after charts without control groups cannot attribute citation changes to content decisions.
- Play 1: Freeze the metric - define citation rate with a fixed denominator, numerator, and exclusions before testing.
- Play 2: Match controls - select pages with similar pre-test citation visibility and content type, not random ones.
- Play 3: Pre-register - write and date a document that locks in your design, timeline, and analysis plan.
- Play 4: Collect data - run queries with a scripted method and log citations in a structured, auditable format.
- Play 5: Label evidence - assign a Level 1, 2, or 3 class and state what your design can and cannot prove.
Key takeaways
Freeze your estimand, denominator, and exclusion rules in a prespecification document before touching any content.
Match control pages on pre-test citation rate and content type - not domain authority or random selection.
Set the observation window length based on your topic's query volatility and known crawl delays, not a universal rule.
Use difference-in-differences to subtract background citation movement from your treatment effect.
Label every result with an evidence class (Level 1, 2, or 3) and state what the design cannot prove.
A single controlled observation is one data point; accumulate multiple tests before drawing directional conclusions.
o choose a control group for an AI citation content test, freeze your estimand first - the exact metric you want to attribute to the content change - then match control pages on pre-existing citation visibility before you touch any content. Without that sequence, your before-and-after chart is noise dressed as evidence. The rest of this playbook is the protocol.
The core failure mode is familiar: an operator updates a set of product pages with better sourcing, runs the same queries some weeks later, sees more citations, and calls it a win. A simple time series of your own pages cannot separate those forces. You need a matched control group that experienced the same external environment while receiving no treatment.
Denominator table: what changes the count at each level
| Unit of observation | Numerator | Denominator | Common exclusions |
|---|---|---|---|
| Content template (e.g., all comparison pages) | Queries where any URL in the template group is cited | Total queries × number of URLs in the group | Queries that resolve to a different template type, duplicate-URL appearances in one answer |
| Domain | Queries where any domain URL is cited | Total queries in the fixed query set | Queries where the domain appears only in a non-citation context (e.g., inline mention without a source link) |
Before you select a single control page, write down the unit of observation, the numerator definition, the denominator definition, and every exclusion rule. This is your prespecification document. Its purpose is to prevent you from adjusting any of these definitions after you see the data.
Start with the unit: are you measuring a single URL, a content template, or a domain? Each choice produces a different denominator. Next, build a fixed query set of non-branded informational queries that your content could plausibly answer. The size of that set is a design decision you make based on the topic coverage of your test pages - there is no universal number. Exclude brand-navigation queries where you already dominate and queries that return no AI-generated answer at all. Each query in the set is one opportunity, so your denominator is the count of queries multiplied by the number of pages if you test multiple pages at the URL level.
Define what counts as a citation: an exact URL match in the AI answer's source list, or a brand mention with a link? For content tests, exact URL match is the cleaner signal because brand mentions vary by provider and session state. Document the session conditions - signed-out, no personalization, consistent locale - and record them in the prespecification document.
A/B Testing 101 - NN/G states: "A/B testing can help UX teams determine the improvements in the user experience that are best for their business goals ."
Operator checkpoint
Your metric is citation rate: citations to a target URL divided by opportunities (queries) tracked. Lock the formula before you run a single query.
You need matched pairs: for each page you intend to change (treatment), select a control page that had a similar citation rate in the baseline period, covers a comparable subtopic, and was published in a similar timeframe.
Run your fixed query set against both groups before any content change and record the citation count for each page. Calculate the average citation rate and its distribution for each group. If the control group's average rate is materially higher than the treatment group's, your post-test comparison will be biased upward - you will understate the treatment effect. If it is materially lower, you will overstate it. Rebalance the groups until the pre-treatment rates are close. If you cannot find adequate matches within your own site, pause the test and expand your topic coverage first rather than proceeding with a mismatched control.
Record the matching criteria in your prespecification document: which pages are in each group, why they were selected, and what the pre-treatment citation rates were. This record is what makes the design auditable.
Matching trap
Matching on domain authority alone ignores the page-level citation rate. A high-DA domain with zero pre-test citations is not a valid control for a page that was already cited regularly.
See whether EdenRank fits the workflow
Request access to the Citation OS and review the process, evidence boundaries, and plan fit. Browse all free tools
Pre-registration is the single step that separates a credible controlled observation from a post-hoc rationalization. Write every parameter of your test in a document, date it, and share it with at least one other person before you make any content change. The document must include: the treatment description (exactly what you will change and on which pages), the control group list, the fixed query set, the providers you will query, the primary metric formula, the exclusion rules, and the planned analysis method.
The timeline structure matters. After the baseline period, assign control and treatment units formally (record the date). Make the content change. Then wait a post-treatment observation period before running the post-intervention queries.
You set this window based on your knowledge of the providers you are testing; there is no universal duration that applies across all providers and topics.
After the observation window closes, run the analysis exactly as prespecified. Do not add queries, swap pages, or change the citation definition. If you discover a data-collection error, document it and report it as a limitation rather than silently correcting it.
- Write the prespecification document with all parameters listed above and record the date
- Run the fixed query set on both groups and record the baseline citation count per page
- Verify that treatment and control groups have comparable pre-treatment citation rates; rebalance if not
- Run the fixed query set on both groups again using the same session conditions as baseline
- Calculate the difference-in-differences: (treatment post − treatment pre) − (control post − control pre)
Checklist
- Confirm that make the content change on treatment pages only; record the exact date of each change
- Confirm that wait the pre-specified post-treatment observation window before collecting post-intervention data
- Confirm that report the result with the evidence class label and the limitations stated in the prespecification document
Operator checkpoint
Pre-registration prevents retroactive adjustment of the hypothesis, control group composition, or success metric after seeing results. A dated document is the cheapest insurance you have.
Executing the measurement requires a scripted approach to avoid manual inconsistencies. Use the exact same prompt format and session state for every query in every collection period.
A screenshot folder is not a structured log - it cannot be queried or audited. If you automate collection, record the script version and any API parameters alongside the data so the collection method is reproducible.
Calculate citation rate for each group and period: citations divided by opportunities. Then calculate the difference-in-differences. A synthetic illustration (labeled as such): if treatment pages gained eight citations and control pages gained three citations over the same period, the adjusted treatment effect is five - not eight. The control group corrected for three citations that appeared regardless of your content change. Real data will not be this clean; document the actual numbers and their variance.
Log field minimum
Every row in your citation log needs: query text, provider, model version, collection date, target URL, citation present (yes/no), session state, and anomaly flag. Missing any field makes the record unauditable.
Assign an evidence class before you share the result. The class communicates what the design can support and what it cannot. Three levels cover most content test designs: Level 1 is an uncontrolled observation - a before-and-after on a single group with no control. It carries no causal weight and should be labeled as such. Level 2 is a controlled observation - a matched control group with a pre-post design and a difference-in-differences calculation. It suggests attribution but cannot rule out selection bias or unmeasured confounders.
Level 3 is a randomized experiment - random assignment of content to treatment and control with a holdout, which is rare in content testing but produces the strongest attributable evidence.
Prespecify the evidence class you are aiming for before the test. That is a defensible result for a skeptical stakeholder. What you cannot do is run a Level 1 design and report it as if it were Level 2.
Your pre-registration document should state the evidence class, the known limitations, and the confounders the design cannot exclude. When you share the result, lead with the evidence class label before the number.
Level 1 vs Level 2
Before
Level 1 - Uncontrolled observation: single group, before-and-after only. Cannot separate content effect from external changes. Report as: 'Citation count changed; cause unknown.'
After
Level 2 - Controlled observation: matched control group, pre-post design, difference-in-differences. Suggests attribution; cannot rule out selection bias. Report as: 'Adjusted treatment effect was X citations in a controlled observation with stated limitations.'
References and further reading
These links are provided for direct inspection. A reference is not treated as proof of every statement in this article.
- 1.A/B Testing 101 - Nielsen Norman Groupnngroup.com
- 2.Using your data with Azure OpenAI - Microsoft Learnlearn.microsoft.com
- 3.Search Engine Land: What 25,000 URLs Reveal About AI Citationssearchengineland.com
- 4.
- 5.Search Engine Journal guide to earning AI and LLM recommendationssearchenginejournal.com
- 6.Cloudflare Radar: AI Search Crawl-to-Refer Ratioblog.cloudflare.com
- 7.RFC 9309: Robots Exclusion Protocolrfc-editor.org
- 8.RFC 9111: HTTP Cachingrfc-editor.org
Written by
EdenRank Editorial Team
The product and editorial team documents repeatable ways to inspect AI-answer visibility, source evidence, and content operations.
Expertise
Want insights like this for your own brand?
Talk to the teamKeep building the topical graph.
Observed Change vs Causal Lift: How to Label AI Citation Evidence
A citation increase after a page change is an observed change. Call it causal lift only when assignment, controls, timing, and frozen evidence support that claim.
How to Map Cited Source URLs Before Planning Content Distribution
Do not start with a list of places to post. Start with the URLs answer engines already cite, classify their ownership and role, then choose the next reachable source lane.
How to Configure robots.txt for AI Crawlers in 2026-Without Guessing
robots.txt is a published crawl preference, not an authentication or security boundary. Configure explicit groups, test real URLs, and verify behavior in logs.
Related AI answers
- How can I use AI visibility data to prioritize which content to update for better AI citation rates?
- What is the best AI citation tracking and visibility tool for a SaaS brand in 2026?
- How can I compare my brand's AI citation frequency across different AI assistants like ChatGPT, Gemini, and Perplexity?