Eight people are recruited, given tasks, and watched. Seven complete every task, the debrief is full of phrases like "clean" and "easy to follow", and the team concludes the user experience is solid. Six weeks later the page experience signals still look poor and nobody can reconcile the two findings.
They are not in conflict. A usability test and a field measurement ask different questions, of different people, under conditions that barely overlap.
What does a usability test measure that a metric cannot?
Comprehension. Whether a person understands what the page is offering, trusts it, and can find the thing they came for — none of which any automated signal observes.
A page can load instantly, hold its layout perfectly, respond to every tap without delay, and still fail every visitor because the pricing is ambiguous or the form asks for a phone number before it says what happens next. No measurement catches that. It is visible in ninety seconds of watching somebody hesitate, and it is invisible in every dashboard you own.
Why does a well-liked site fail the measurements anyway?
Because the measurements sample conditions your test never included. Field data is collected from whatever devices and networks real visitors bring; a usability session is typically run on good hardware, on a fast connection, with the page already warm.
Consider what differs between the two populations. Your participants were told which site they were visiting, so they waited through the load. A real visitor arriving from a search result has no such commitment and leaves during the same delay. Your participants used the device you provided. Real traffic arrives on a spread of phones, some of them several years old, on connections that vary by the minute.
Usability testing samples people carefully and conditions badly. Field measurement samples conditions perfectly and understands people not at all. Each is blind exactly where the other sees, which is why a site can pass one convincingly while failing the other.
Which problems does each method catch?
Sorting them explicitly is usually enough to end the argument about which report to believe.
| Problem | Usability test | Field measurement |
|---|---|---|
| Confusing navigation labels | Catches it | Invisible |
| Slow load on a mid-range phone | Usually misses it | Catches it |
| Form that asks too much too early | Catches it | Invisible |
| Layout that shifts as ads load | Sometimes catches it | Catches it |
| Unclear pricing or scope | Catches it | Invisible |
| Tap target that misfires on small screens | Catches it if tested on a phone | Partially |
| Interaction delay from a heavy script | Rarely reproduced | Catches it |
The pattern is that anything about meaning belongs to the test and anything about cost belongs to the measurement. A team that reads both as one score will keep finding contradictions that are simply the two halves of the picture.
Does passing the measured signals mean the experience is good?
No. Passing means nothing about the delivery of the page is actively driving visitors away. It is a floor, not a verdict on quality.
This matters for how the work is prioritized. A site that already passes gains little from further performance work and may have a substantial problem in what the pages actually say — which is the territory that E-E-A-T concerns describe and no speed metric will ever surface. Continuing to optimize the number that already passes is how a team stays busy while the real problem sits untouched.
Running the usability test on a laptop for a site where most visits are mobile. It produces confident findings about an experience almost none of your visitors have, and it hides every problem that only appears on a small screen.
Where do the two methods disagree most?
On pages carrying third-party content. Embeds, chat widgets, review carousels and tag-managed scripts are close to invisible in a moderated session and dominate what field measurement records.
They also disagree about anything that only appears under load or on a cold cache. A participant who visits three times during a session gets a warm cache on visits two and three; the majority of your real traffic is on visit one, forever. Testing the second impression of a page and measuring the first is one of the quieter ways these two methods end up describing different websites.
How do you run both without doubling the work?
Let each method scope the other. Use the field data to choose which pages to test and the test to explain what the field data cannot.
In practice that means starting from the pages where measured experience is worst and organic search traffic is highest, then testing those specific pages on a real mid-range phone rather than on the machine sitting on the desk. Five sessions on the two pages that matter beat twelve on a representative sample, and the participants will surface the comprehension problems that no amount of instrumentation would have shown you.
Look at which of your last three UX decisions came from watching somebody and which came from a dashboard. If they all came from one source, you have been fixing half of the experience and measuring the other half.
Site tests well and still underperforms in search?
We will separate what your visitors struggle with from what the measured signals penalize, and tell you which one is costing you more.
Get in Touch →Does this apply to your site?
Reading about it is one thing. Point the scan at your own site and see whether this applies to you, and what it is worth fixing.
Free and unlimited. No account, no card, and you get every finding rather than a teaser.