An entire industry sells thumbnail optimisation. We measured 440 of them across two and a half years to find out whether the things those tools measure actually predict anything. They mostly don't.
We took one niche — “AI tools” — and sampled the 40 most-viewed videos per quarter for eleven consecutive quarters, from the start of 2024 to the third quarter of 2026. Only standard-length videos (4–20 minutes) were included, so every thumbnail is the same 16:9 format and Shorts don't contaminate the sample. That gives 440 thumbnails, evenly spread across time.
Each thumbnail was downloaded at full resolution and measured on seven properties, computed from the raw pixels:
Then we asked two questions. First: do these properties drift over time? Second, and more usefully: do they predict how well a video does?
To test prediction we standardised every metric within each quarter, so era effects can't leak in, and correlated it against performance. Here is every correlation we measured, against views adjusted for channel size:
Against raw view counts, with no adjustment at all, the picture is
even flatter — the largest correlation of the seven is r = 0.09, which
is roughly 0.8% of the variance. Brightness, saturation, colourfulness,
contrast, busyness: none of them move the needle.
| Metric | r (vs views) | r (vs views/subs) | Verdict |
|---|
Colourfulness came back at r = −0.152 (p = 0.0014) —
comfortably significant, and significant even after correcting for seven tests.
Read naively it says something genuinely interesting: less colourful thumbnails
overperform, the exact opposite of standard YouTube advice.
It is not true. The outcome variable — views per subscriber — is mechanically lower for large channels, because a channel with 10 million subscribers cannot get 50 views per subscriber. So any thumbnail property that correlates with channel size will produce a fake correlation with that outcome. We checked:
Colourfulness correlates with channel size at r = +0.133
(t = 2.80, p = 0.005). Bigger channels make more
colourful thumbnails — they have designers, brand systems and budgets.
Bigger channels also mechanically have lower views-per-subscriber. That chain
alone predicts a spurious negative correlation of roughly the size we observed.
Once channel size is accounted for, the effect is not distinguishable from zero. What looked like a thumbnail finding was a production-budget finding.
This is the trap the whole category falls into, and it's worth stating plainly: polished thumbnails correlate with success because successful channels can afford polish, not because polish causes success. Correlational thumbnail advice almost never separates those two.
The second question was whether the look of a niche visibly drifts. Each panel
below is one metric, quarter by quarter, on its own scale. r² is how
much of the movement a straight-line trend explains.
So the honest answer to “is this niche's visual style evolving?” is: if it is, the effect is smaller than the noise between individual thumbnails in any single quarter. Within one quarter the spread on every metric is wider than the entire 2.5-year range of the quarterly averages.
We are reporting a null result, and null results are easy to overclaim. So, precisely:
It tells you that these seven pixel statistics, in this one niche, don't predict performance. That is a much narrower claim.
The specific weaknesses, stated openly:
What we'd claim with confidence is narrower and still useful: if a tool scores your thumbnail on brightness, saturation, contrast or colourfulness and tells you that score predicts performance, it is not supported by this data. Those numbers are descriptive, not predictive, and should be sold as such.
Paste any YouTube URL or video ID. Your thumbnail is measured on the same seven properties and placed against the distribution of all 440 thumbnails in the study. Nothing is uploaded — the analysis happens on your device.
Thumbnails fetched at maxresdefault (falling back to
hqdefault), downscaled to 160×90 and measured in a canvas context.
Colourfulness follows Hasler & Süsstrunk (2003). Skin detection uses the
standard YCbCr bounding rule. Edge density is the mean Sobel gradient magnitude
over the luminance plane. Video metadata — publish date, views, subscriber count —
came from the vidIQ API, sampled by quarter and ordered by view count.
All correlations are Pearson on within-quarter z-scores; significance is a
two-tailed t-test with 438 degrees of freedom; the multiple-comparison threshold
is Bonferroni at α = 0.05/7 = 0.0071. The analysis pipeline is the same code
running in the tool above — you can read it by viewing the source of this page.
The full run was executed twice from a cold start and reproduced identically to the displayed precision.