Verdict

Verdict Research

We graded the reviewers: the same product swings 0.6 points depending on who reviews it.

By Mike Hunter · June 16, 2026 · Methodology at the bottom

We aggregate professional review scores for a living, which means we can see the same product through dozens of reviewers' eyes at once. So we asked: do review outlets actually grade on the same curve? We measured the scoring bias of 31 publications across 1,447 products and more than 5,000 individual expert scores. They don't. The outlet you happen to read can move a product's score by more than half a point — enough to flip a “great” into a “just okay.”

0.6 pt
Gap between the toughest and most generous outlet, after controlling for which products each one reviews.
22%
Of products where two outlets disagree by a full point or more on the same item.
31
Review outlets ranked, across 1,447 products and 5,000+ scores.
Diverging bar chart of 12 review outlets ranked by how far their scores sit from the cross-outlet consensus. OutdoorGearLab grades 0.37 points below consensus (toughest); CleverHiker grades 0.24 above (most generous).
Each outlet's average distance from the per-product consensus, across 1,447 products. Negative = grades tougher than its peers.

Reviewers don't grade on the same curve

To compare outlets fairly, we can't just rank them by average score — an outlet that mostly reviews premium flagships would look artificially generous. Instead, for every product we compute the consensus (the average score across all outlets that reviewed it), then measure how far each outlet lands above or below that consensus on the products it covers. That difference is its grading bias. Across the 31 outlets with at least 20 products in our database, the spread runs from 0.37 points below consensus to 0.24 above — a 0.6-point swing on a 5-point scale that depends entirely on who's holding the pen.

The toughest and most generous graders

Bias is each outlet's average distance from the per-product consensus — negative means it scores tougher than its peers, positive means easier. “Avg” is its raw average score on a 5-point scale; “n” is how many products it scored.

Grade hardest

  • OutdoorGearLab-0.37 · 3.78 avg · n=97
  • SoundGuys-0.34 · 3.87 avg · n=44
  • Wired-0.23 · 3.93 avg · n=23
  • Tom's Hardware-0.16 · 4.01 avg · n=73
  • TechGearLab-0.16 · 3.99 avg · n=101
  • PCMag-0.15 · 3.98 avg · n=130

Grade easiest

  • CleverHiker+0.24 · 4.5 avg · n=44
  • What Hi-Fi+0.23 · 4.55 avg · n=39
  • Serious Eats+0.16 · 4.58 avg · n=20
  • Trusted Reviews+0.12 · 4.33 avg · n=36
  • Digital Camera World+0.11 · 4.47 avg · n=51
  • TechRadar+0.06 · 4.27 avg · n=191

The same product, a full grade apart

On 22% of products, two outlets we track disagree by a full point or more — and these aren't obscure picks. Here are clean splits between well-known publications on products you've probably shopped for:

ProductLiked itLess sold
Bowers & Wilkins Px8
Noise-cancelling headphones
What Hi-Fi
5/5
PCMag
3.5/5
Bose QuietComfort Earbuds Ultra
Wireless earbuds
What Hi-Fi
5/5
SoundGuys
3.7/5
Sabrent Rocket Nano V2
Portable SSDs
PCWorld
4.5/5
Tom's Hardware
3/5
Eero 7
Mesh Wi-Fi systems
TechGearLab
4.4/5
PCMag
3/5
Razer Kiyo Pro Ultra
Streaming webcams
PCWorld
4.5/5
Wired
3/5
JBL Xtreme 4
Outdoor Bluetooth speakers
What Hi-Fi
5/5
SoundGuys
3.65/5

Why this happens — and why aggregation fixes it

None of these outlets is “wrong.” Different test benches, house styles, and audiences produce different scores: a headphones specialist weighs sound the way a general-tech site never would, and some publications simply grade on a stricter curve. The problem is for the reader, because a single review — generous or harsh — can swing a buying decision. Averaging across outlets cancels out each one's individual bias, which is the entire premise of Verdict.

What this means if you're buying

Don't let one review — high or low — decide it for you. Check whether the outlet you're reading tends to grade soft or hard, and look at where a product lands across several reviewers. Our category rankings do exactly that — every score is an average across the publications that tested it. See also our companion study on whether price predicts quality (it barely does).

Methodology & data

Figures come from Verdict's database as of June 2026: 5,000+ individual review scores normalized to a 5-point scale, drawn from professional publications across 295 categories. For each product with two or more outlet scores we compute a consensus (the mean across its outlets); an outlet's grading bias is the average of (its score − that product's consensus) over every product it scored. The leaderboard includes the 31 outlets with 20 or more scored products; the 0.6-point figure is the range from the lowest to highest bias among them. “Disagree by a full point” counts products (with three or more outlet scores) whose highest and lowest scores differ by 1.0 or more. Named per-product splits use only recognizable publications with both scores between 3 and 5; low-end score-extraction outliers were excluded. Full ranking methodology: Verdict methodology. Questions or want the underlying numbers? Get in touch.