Setting a golden standard to run GEO experiments for the marketing community. This document explains how.
A Note on How We Write About Research
Disclaimer: OtterlyAI follows the structured research methodology of Kohavi et al, which is described in this document. However, our primary audience is marketing professionals, not statisticians. Our published experiment articles are written as journalistic pieces: what matters most to marketers (the findings, the practical implications) comes first. Technical methodology details sit in a supporting role.
That means you will find business and marketing takeaways at the top of every OtterlyAI experiment article, before you see methodology or caveats. That is intentional. The methodology is always there, always consistent, and always linked back to this page.
Another caveat is that AI Search is not randomized, it never will be. AI Search Engines already pre-select and pre-filter sources for their users, meaning we cannot follow Kohavi et al‘s methodology to the letter as that would require a 100% randomized population. Our methodology still follows the principles as much as possible set by R. Kohavi and his team. Below is our framework to inspire future researchers.
Explore All GEO Experiments & AI Search Case Studies
You can download the latest version of the experiment sheet below. It gets updated every week with new experiments as we test them, so bookmark this page or check back regularly to follow our progress.
👉 View GEO Experiment/Studies (Google Sheets)
Key Takeaways
- OtterlyAI experiments change only one variable at a time. Everything else stays identical across versions.
- Every experiment has at least one control group: a version where nothing changes, used as the baseline.
- The success metric (OEC) is defined before the experiment starts, not after results come in.
- Results must be practically meaningful to be acted on, not just statistically detectable.
- All analysis decisions are documented before data collection begins.
- OtterlyAI has its own research methodology but it is heavily inspired by Kohavi et al.
Why This Document Exists
GEO (Generative Engine Optimization) is a new field. There are no widely agreed standards for how AI search experiments should be designed or reported. That creates a problem: results from poorly designed studies get published, circulated, and acted on even when the methodology does not hold up.
OtterlyAI publishes original research on AI search visibility. This document explains the standards we follow. Every published study links back to this page so readers always have a way to evaluate the rigor behind the findings.
The principles described here are not invented by OtterlyAI. They come from established research design frameworks used in psychology, medicine, engineering, and digital product development. We apply them to the specific context of AI search visibility.
The Core Idea: Change One Thing at a Time
Every OtterlyAI experiment is built on one principle: if you want to know whether something causes a result, you need to make sure only that one thing changes. Everything else stays identical.
This is the logic of a controlled experiment, formalized by statistician Ronald Fisher in 1935 and used as the foundation of experimental research ever since. Where we can randomize, we run controlled experiments. Where the setting does not allow it, for example when we cannot control how an AI Search Platform treats a page, we use quasi-experimental designs that apply the same one-variable logic across matched groups. Both approaches follow the same rule: only one thing changes, so any difference in results can be traced to that one thing.
If two things change at the same time, you cannot tell which one caused the result.
A practical example: in a study testing whether Reddit engagement affects AI citation rates, every version posts the same content on the same day. The only difference is how much engagement each post receives. If citation rates differ between versions, the engagement is the most likely cause. If content or posting time also varied, the result would be uninterpretable.
Glossary: AI Search Experiment Terms Explained
These are the terms you will encounter across OtterlyAI research publications. Each definition is written for marketers, not statisticians.
| Term | What it means |
| Hypothesis | A one-sentence prediction stated before the experiment begins: ‘If X changes, then Y will happen.’ It must be falsifiable, meaning the experiment can prove it wrong. |
| Control (Variant A) | The version of your experiment where nothing changes. It is the baseline you compare everything else to. Think of it as the ‘do nothing’ option. |
| Treatment (Variant B/C/..) | A version where you deliberately change one thing. Also called a ‘variant.’ You can have multiple treatments in one experiment (hence A/B/N). |
| Canary/Watermark method | Embedding a unique, traceable signal (such as a specific number or phrase) into each experiment version so you can identify which exact version an AI Search Platform picked up. Like a fingerprint per page. |
| Pre-registration | Documenting all your experiment decisions before collecting data: what you are testing, how you will measure it, and what counts as a meaningful result. Prevents cherry-picking results after the fact. |
| The success metric or OEC (Overall Evaluation Criterion) | The single metric that defines success for the experiment. Chosen before the experiment starts, not after you see results. (e.g. Total AI Citations) |
| Independent variable | The one thing you change on purpose between versions. Everything else stays the same. |
| Dependent variable | The outcome you measure. In GEO experiments, this is usually whether a page gets cited by an AI Search Platform |
| Dose-response relationship | When higher levels of your treatment produce proportionally stronger results. For example: low engagement produces some citations, medium engagement produces more, high engagement produces the most. A consistent dose-response pattern makes the causal argument much stronger. |
| Internal validity | Confidence that your result was caused by the treatment, not by something else that happened at the same time. |
| Practical significance | Whether a result is large enough to matter, not just whether it is detectable. A tiny improvement might be real but still not worth acting on. |
| Novelty effect | When early results look unusually strong simply because the content or page is new. Fades over time as the newness wears off. |
| Longitudinal study | An experiment that tracks the same thing over an extended period of time, measuring how results change week over week or month over month. |
| Between-group comparison | Comparing two or more separate groups rather than the same group before and after a change. OtterlyAI uses this when running experiments across different domains or subreddits. |
OtterlyAI’s 6 Research Methods at a Glance
| Method | Best for | OtterlyAI example | Source |
| 1. A/B/N Testing | Comparing multiple engagement or content levels simultaneously | Reddit engagement experiment (A = no engagement, B/C/D = increasing levels) | Fisher (1935); Kohavi, Tang & Xu (2020) |
| 2. Canary / Watermark Method | Tracing which specific signal an AI Search Platform picked up when multiple signals exist on a page | Image naming experiment: each page used a unique number so the AI answer revealed which signal it read | Used in security research and large-scale web crawl analysis |
| 3. Single-Variable Controlled Domain Test | Testing one content or technical change across two matched domains with identical structure | Capybara domain experiments: author attribution vs no author, pillar page vs split articles | Montgomery (2017); Campbell & Stanley (1963) |
| 4. Platform-by-Platform Observation | Measuring whether a change produces consistent or fragmented results across different AI platforms | Schema markup experiment: some platforms could read schema, others could not | Shadish, Cook & Campbell (2002) |
| 5. Longitudinal Observation | Tracking how a single change evolves over time, particularly for signals with delayed effects | llms.txt 90-day study: measuring AI bot visits to /llms.txt over three months | Hill (1965) |
| 6. Between-Group Comparison | Comparing two separate communities or environments with identical content but different context | Reddit active vs dormant community: same posts in r/GEOexperiments vs r/GEOexperts | Shadish, Cook & Campbell (2002) |
Method 1: A/B/N Testing
Most people are familiar with A/B testing: run two versions of something, compare the results, see which performs better. A/B/N testing extends this to three or more versions. The N simply means any number of variants.
OtterlyAI uses A/B/N testing when we want to understand whether more of something produces proportionally more of a result. This is called a dose-response relationship, a concept used widely in medical research (Hill, 1965) and adapted here for digital experimentation.
| Version | Role | What changes | What stays the same |
| A | Control | Nothing. Baseline only. | Content, timing, format, platform |
| B | Treatment 1 | One thing, at a low level | Content, timing, format, platform |
| C | Treatment 2 | Same thing, at a medium level | Content, timing, format, platform |
| D | Treatment 3 | Same thing, at a high level | Content, timing, format, platform |
If citation rates increase consistently from A to B to C to D, that is a much stronger finding than a single A vs B comparison. It shows the effect grows with the treatment, not that it appeared once by chance.
OtterlyAI example: Reddit engagement experiment. Four subreddits post identical content daily. Version A receives no engagement. Versions B, C, and D receive increasing levels of upvotes and comments. The question: does higher engagement lead to more AI citations?
Method 2: The Canary / Watermark Method
The canary method is used when a page has multiple signals and you want to know which specific one an AI Search Platform picked up.
The idea comes from canary deployments in software engineering, where a unique identifier is embedded so you can trace exactly which version of a system produced a given output.
In OtterlyAI experiments, we apply this by embedding a unique, traceable piece of information into each page variant. Each page gets a different number or phrase. When an AI Search Platform answers a query about that topic, the specific detail in its answer reveals which page signal it used.
Think of it as a fingerprint per page. The AI’s answer tells you which fingerprint it read.
OtterlyAI example: Image naming experiment. Each page embedded a different number (e.g. ‘3 vegetarians in the OtterlyAI team’ on one page, ‘7 vegetarians’ on another). Each page placed the number in a different location: image filename, alt text, figcaption, body text, or schema markup. When AI platforms answered ‘how many vegetarians are on the OtterlyAI team?’ the number in the answer revealed exactly which signal the AI read.
Method 3: Single-Variable Controlled Domain Test
This method runs two separate websites (domains) in parallel. Both sites cover the same topic, use the same structure, the same word count, and the same publishing cadence. Only one variable differs between them.
This approach is drawn from factorial experimental design (Montgomery, 2017), adapted for web content. It is particularly useful for testing content and technical signals that cannot easily be isolated within a single site.
OtterlyAI example: Author attribution experiment. Two test domains (capybara-world.com and capybaraverse.com) publish matched content. One domain displays author bios; the other publishes without authorship. Citation rates are tracked across both to measure whether author attribution affects AI visibility.
A key rule for this method: only one variable may differ. If one domain also has a different URL structure, different word count, or different internal linking, the result becomes uninterpretable.
Method 4: Platform-by-Platform Observation
Not all AI Search Platforms behave the same way. A technical change that works on Google AI Overviews may have no effect on Perplexity. OtterlyAI runs observations across all six monitored platforms simultaneously: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini, and Microsoft Copilot.
When results differ by platform, the finding is reported platform by platform. A consistent result across all six platforms is a stronger finding than a result observed on one.
OtterlyAI example: Schema markup experiment. Schema was added to test pages and monitored across all six platforms. Some platforms showed an increase in citations after schema implementation; others showed no change or a decrease. The fragmented result was reported honestly, per platform, with no overall claim made.
Method 5: Longitudinal Observation
Some signals need time to show an effect. A new file added to a website may not be crawled for weeks. A new Reddit community may take months before AI Search Platforms begin citing it consistently.
Longitudinal observation tracks the same metric over an extended period, measuring whether and when a change produces an effect. The method is grounded in Bradford Hill’s (1965) criteria for causal inference, which require consistency of effect over time as one indicator of a real relationship.
OtterlyAI example: llms.txt 90-day study. A /llms.txt file was added to the OtterlyAI website. AI bot traffic to that file was tracked every week for 90 days. Under 0.1% of total AI bot visits went to /llms.txt, leading to the finding that llms.txt is not a meaningful AI crawl signal.
Method 6: Between-Group Comparison
A between-group comparison tests whether two separate environments with identical content produce different outcomes based on one contextual difference.
Unlike the controlled domain test, which isolates a content or technical signal, between-group comparisons isolate a contextual signal: the environment the content lives in. This method follows the quasi-experimental design principles described by Shadish, Cook, and Campbell (2002).
OtterlyAI example: Reddit active vs dormant community experiment. Two subreddits (r/GEOexperiments and r/GEOexperts) publish the same content. One subreddit has active community engagement; the other does not. The question: does community activity level, independent of content, affect AI citation rates?
The Rules Every OtterlyAI Experiment Follows
Regardless of which method is used, every OtterlyAI experiment follows these rules. They are not guidelines. If an experiment cannot meet these standards, it is not published as a controlled or quasi-experimental study.
| Rule | What it means in practice |
| One variable at a time | Only one thing changes between versions. If two things change, you cannot tell which one caused the result. |
| Define success before you start | The primary metric (OEC) is locked before data collection begins. It does not change after results come in. |
| Run all versions at the same time | Parallel execution removes the effect of timing. A Google update or viral Reddit post hits all versions equally. |
| Lock the experiment once it starts | No changes to format, posting time, or treatment after Day 1. Mid-run changes invalidate the data. |
| Document everything before you look at results | All analysis decisions are pre-registered. Any decision made after seeing data is labeled exploratory. |
| Set a minimum effect size | Results must clear a practical threshold to be acted on. Small, detectable differences are reported but not treated as recommendations. |
How to Read an OtterlyAI Experiment Article
Every OtterlyAI experiment article is structured the same way. Marketing findings come first. Methodology comes after. This is intentional: the article is written for practitioners, not researchers.
Here is what each section contains:
- Key Findings: the two or three results that matter most from a marketing perspective. Read this if you want the short version.
- Hypothesis: the question the experiment was designed to answer, stated before the experiment began.
- Scope of Study: what was tested, on which platforms, over what time period.
- Methodology: which method was used, how versions were structured, and what was kept constant.
- Results: the full findings broken down by version and platform.
- What This Means: practical implications for GEO strategy.
- Limitations: what this experiment does not prove, and what future research would need to address.
The limitations section is not a disclaimer added to cover weaknesses. It is a required part of honest research reporting. Readers should treat any study that does not include limitations with more skepticism, not less.
What OtterlyAI Research Cannot Claim
AI Search Platforms do not publish their ranking or citation logic. OtterlyAI measures what users see in AI-generated answers via public web interfaces. We do not have access to the internal workings of any AI platform.
This means every OtterlyAI finding is an observed association, not a guaranteed mechanism. We report what we measured, across the prompts and platforms we tracked, during the time period the experiment ran.
- A finding on ChatGPT does not automatically apply to Perplexity or Gemini.
- A finding from one prompt set does not automatically generalize to all prompts in a category.
- AI platform behavior can change at any time without notice. Findings are dated and reflect behavior at the time of the experiment.
- Observed correlations do not prove platform-level causation. We use dose-response patterns and controlled conditions to build the strongest possible causal argument, but we do not claim certainty.
Got a GEO experiment idea? Let’s collaborate
At OtterlyAI, we don’t write about AI Search from the sidelines. We test it. With the community, we run and update a public GEO Experimentation Tracker showing what we are running and what moved AI Search. If you are running your own GEO experiment, email rick.tousseyn@otterly.ai and we may feature it with your name attributed.
👉 Want to measure your AI Visibility? Start with OtterlyAI and our GEO Guide.
👉 Want to check if you are accidentally blocking AI crawlers like Gartner is? Run your domain through the OtterlyAI Crawlability Checker inside the GEO Audit.
References
The methodological frameworks described in this document are grounded in the following published works.
- Kohavi, R., Tang, D. and Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.
- Campbell, D.T. and Stanley, J.C. (1963). Experimental and Quasi-Experimental Designs for Research. Houghton Mifflin.
- Fisher, R.A. (1935). The Design of Experiments. Oliver and Boyd.
- Hill, A.B. (1965). The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine, 58(5), 295-300.
- Montgomery, D.C. (2017). Design and Analysis of Experiments (9th ed.). Wiley.
- Shadish, W.R., Cook, T.D. and Campbell, D.T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.



