Somebody proposes a change to every title tag on the site. Or a new template, or a schema rollout, or removing a block from two thousand pages. The change is deployed everywhere at once, results are ambiguous for six weeks, and by the time anyone is sure, three other things have changed too.
Search work cannot be split-tested the way conversion work can — you cannot show Google two versions of a page and compare. You can still stage a rollout so a bad change costs a subset instead of the site.
Why not a normal A/B test?
Because the unit of measurement is a URL, not a visitor. Serving different content to different visitors at the same URL does not create two ranking outcomes; it creates one, and if the variation is served based on who is asking, it edges toward cloaking.
The workable design splits URLs rather than users: apply the change to one group of pages, leave a comparable group unchanged, and compare the groups over time.
The control group is what makes this a test. Without one you are measuring the change plus seasonality plus whatever Google did that month, and attributing all of it to your edit.
How do you pick the two groups?
From pages that behave alike. The comparison only works if the groups would have moved together without your change, so they need similar templates, similar intent and similar traffic levels.
Practical selection:
- Take one template with enough pages to see a pattern — thirty is workable, a hundred is better.
- Sort by impressions and alternate down the list into two groups, so both contain a similar mix of strong and weak pages.
- Exclude anything seasonal, recently changed, or unusually volatile.
- Record both lists before you start. Deciding afterwards which pages were in the test is how a null result becomes a positive one.
How long before you can read it?
Longer than feels reasonable. Google has to recrawl and re-evaluate, and the effect appears gradually. Four weeks is the earliest most changes say anything; six to eight is more honest, and title changes tend to show faster than structural ones.
Two rules make the wait worthwhile. Change nothing else on those pages during the window, and compare the same weekday ranges, because weekly cycles will otherwise supply whatever result you were hoping for.
Calling the result after ten days because the line moved. Ten days of search data is mostly noise, and the direction it happens to point is not information.
What do you measure?
Depends on the change, and picking the metric afterwards is how teams talk themselves into results.
| Change | Metric |
|---|---|
| Title or description rewrite | Click-through rate at stable positions |
| Content expansion | Impressions and number of ranking queries |
| Internal linking | Impressions on the target pages |
| Schema deployment | Rich result impressions in Search Console |
| Performance work | Field data at the 75th percentile |
Write down the metric and the expected direction before deploying. It is a small discipline that removes the most common failure in this kind of work.
When should you skip the test?
When the change is a fix rather than a bet. A broken canonical, a missing redirect, a noindex on a page that should be indexed — these are defects, and testing whether fixing a defect helps is a waste of six weeks.
Test the things that are genuinely arguable: title formats, content length, how aggressively to consolidate, whether a block earns its place. Those are the changes where confident opinions are cheap and evidence is not.
One caveat worth stating plainly: this design detects effects large enough to show above the noise, and it will not resolve small ones. If a change is expected to move things by a couple of percent, a thirty-page test will not prove it either way, and the honest response is to decide on reasoning rather than to pretend the data settled it.
Pick the next sitewide change on your list and ask what the control group would be. If you cannot name one, that is the part of the plan still missing.
About to change something across every page?
We will design the rollout so you find out whether it worked before it reaches the whole site.
Get in Touch →Does this apply to your site?
Reading about it is one thing. Point the scan at your own site and see whether this applies to you, and what it is worth fixing.
Free and unlimited. No account, no card, and you get every finding rather than a teaser.