Across 5,000+ professional review scores spanning 285 categories, the price-quality correlation is r = 0.05. Here's the full breakdown.
Every shopper assumes it on some level: spend more, get better. I run a review-aggregation site, which means I had the data to actually test that assumption. It mostly falls apart.
The data and the test
I pulled every product in our database that has both a tracked price and an aggregated review score: 1,434 products, backed by more than 5,000 individual professional review scores from over 1,800 publications, spanning 285 product categories. Scores are normalized to a 5-point scale and averaged per product, so no single outlet's grading habits dominate.
The test is the simplest one possible: Pearson correlation between what a product costs and how well it reviews.
Price and quality are nearly uncorrelated
The correlation is r = 0.05. Statistically, that's noise. Knowing a product's price tells you almost nothing about how good independent reviewers found it.
The average product in the data costs $407 and earns 4.35 out of 5. The most expensive tenth of products, everything above roughly $950, reviews no better on average than the rest. As price climbs from $10 to $8,000, the cloud of scores just stays flat.
In 25% of categories, the cheapest pick beats the priciest
In a quarter of the categories analyzed, the least expensive product scored as high or higher than the most expensive one, usually at a fraction of the price. Some clean examples:
- Full-frame mirrorless cameras: the $3,499 Nikon Z8 (4.5/5) out-rates the $7,199 Sony A1 II (4.2/5).
- 4K monitors: the $391 Dell U2723QE (4.5/5) beats the $1,599 Samsung ViewFinity S9 (4.0/5).
- Robot vacuums: the $399 Mova P10 Pro Ultra (4.3/5) beats the $1,599 Dreame X60 Max Ultra (4.0/5).
- Premium mechanical keyboards: the $120 NuPhy Air75 V2 ties the $300 Glorious GMMK Pro at 4.2/5.
To keep these honest, categories only count when they have four or more priced products and the cheaper product costs under 60% of the pricier one, and score outliers were excluded from named examples.
The premium tax
The flip side is a set of flagship-priced products that review below the median product in the whole dataset:
- Beelink GTR9 Pro (AI mini PCs, $3,499): 3.6/5
- Roborock Saros Z70 (premium robot vacuums, $2,399): 3.5/5
- Nami Klima (electric scooters, $2,999): 3.7/5
- Ecovacs GOAT A3000 (robot lawn mowers, $2,500): 3.9/5
Paying flagship money buys you positioning, not a guarantee of performance.
Reviewers don't grade on the same curve
Averaging across publications exposes something else: among outlets with 30+ scores in the data, there is a 0.75-point gap between the toughest and most generous graders on a 5-point scale. OutdoorGearLab averages 3.79 and SoundGuys 3.87, while What Hi-Fi averages 4.54 and CleverHiker 4.50. The same product can swing dramatically depending on who reviewed it, which is worth remembering any time a single review is about to swing your decision. I published a separate analysis of per-outlet grading bias in this companion study.
Why this happens
My read on the mechanics, having watched this data accumulate: price encodes positioning, brand, and launch-window anchoring far more than it encodes engineering. Categories mature fast. Last year's flagship features show up in this year's mid-range at half the price, while review scores track how a product performs against current expectations, not against its price tag.
What actually predicts quality
Not the sticker. The only reliable signal in the data is what independent reviewers concluded, ideally several of them averaged together so any one outlet's harsh or generous streak washes out. That is the entire reason aggregation works as a method.
Methodology and limitations
Figures come from my database as of June 2026. Correlation is Pearson's r on price vs. aggregated rating. Prices are point-in-time retail listings and do change. Per-publication averages include only outlets with 30+ scores. Two honest caveats: products that get professionally reviewed at all skew toward things worth reviewing, so the dataset underrepresents true junk; and aggregated scores can compress differences at the top of a category. Neither changes the headline result: within the universe of products people actually research, price is a poor proxy for quality. The interactive version with the underlying numbers is on my site.