Everybody treats a YouTube thumbnail as the thing that wins you the click. Researchers sat 24 people in front of an eye tracker, showed them eight real YouTube search results, and asked them to pick a video. The thumbnail did dominate their attention, taking 35% of the time they spent looking while the title took 10%.
Then the researchers listened to people talk through their choices and counted which part of each video card drove each decision. The thumbnail drove 26 of those decisions. Only 5 were somebody deciding to watch a video. The other 21 were somebody deciding to skip one.
A video has to get past that filter before anyone will pick it, which is a different job from the one every thumbnail guide is written for.
I make small apps (Repeat Recorder, Paint Vlix, Camera to Clipboard) and film everything myself, so I wanted the rules somebody had actually measured.
The short version: three things matter, and almost none of the famous advice is one of them.
- People use your thumbnail to decide which videos to skip, so one subject they can recognise at phone size beats a clever idea. When someone scans a page of results, the thumbnail gets about 35% of the time they spend looking at a video's card and the title gets about 10%. People do not stare at the thumbnail for longer than anything else. They glance back at it about four times as often, and each glance lasts roughly a fifth of a second.
- Say something different from the thumbnails beside you, because standing out is only worth anything against your neighbours. Across 8,977 headline tests, being more specific raised clicks 5.5% when the competition was vague and cut them 9.9% when it was already specific. A face works the same way, helping only where the other thumbnails do not have one.
- Stop optimising for the click rate, because YouTube does not. Its own thumbnail test never shows you a click rate and picks winners on watch time, its own help pages teach that a falling click rate is usually a sign of success, and its own engineers wrote that the recommendation system cannot see your thumbnail at all.
Where your effort actually pays on a thumbnail

The whole answer, before the evidence behind it.
- Build the thumbnail around one subject and at most a few large words, because a single look lasts about a fifth of a second. The picture and the writing both have to land in that one glance. On a phone the thumbnail is drawn about 168 pixels wide, so somebody has to recognise what the picture is of, and read any words on it, at that size and without leaning in. Every part of a video card gets looked at for the same 176 to 240 milliseconds. The thumbnail wins because people glance back at it about four times as often, not because any single glance lasts longer. A second subject, a crowded background or a full sentence of text all need a second glance, and nobody gives one.
- Make it different from the thumbnails around it rather than louder than them. Standing out has no value on its own, only against competitors, and the value collapses when every neighbour is shouting too. This is why the same face, the same bold colour and the same three-word caption stop working once everybody uses them.
- Promise only what the video delivers, because YouTube grades the thumbnail on what happens after the click. Its own testing tool picks the winning thumbnail on watch time and never reports a click rate to you at all. A thumbnail that wins the click and loses the viewer is the exact failure that tool is built to catch.
So: one clear subject first, different from your neighbours second, honest throughout. Everything below is either the evidence for those three or the reason a famous rule did not make the list.
Your thumbnail gets four times the looks your title does

The best measurement of this used a real YouTube layout. Bitzenbauer and colleagues, in Research in Science Education, 2024, eye-tracked 24 people choosing a video to teach with from eight real search results, with every part of the card visible: thumbnail, title, channel, subscriber count, views, likes, length and upload date.
| Part of the video card | Share of the time people looked | Share of separate looks | How long each look lasted |
|---|---|---|---|
| The thumbnail | 35.4% the only one above 10% | 45.6% nearly half of everything | About a fifth of a second |
| The title | 9.6% | 11.6% | About a fifth of a second |
| The channel name | 6.2% | 7.1% | About a fifth of a second |
| Subscriber count | 2.6% | 3.0% | About a fifth of a second |
| View count | 2.5% | 3.1% | About a fifth of a second |
| The video length | 1.1% | 1.1% | About a fifth of a second |
- The thumbnail took 35.4% of all the time people spent looking and the title took 9.6%, a gap of roughly four to one. By number of separate looks the gap was similar, 45.6% against 11.6%. The thumbnail was the only part of the card that got more than a tenth of the attention.
- The thumbnail does not hold the eye longer, it simply catches the eye more often. Every single element averaged between 176 and 240 milliseconds per look, and the middle of that spread sat between 200 and 300 milliseconds. So the thumbnail is not more absorbing than the title. It is just returned to four times as often, in glances a fifth of a second long.
- An earlier eye-tracking study found the same roughly two to one gap when people were casually browsing rather than searching. Yangandul and colleagues reported at a 2018 eye-tracking conference that people gave almost twice as much attention to thumbnails as to title text, and looked at the thumbnail first. Their task was casual browsing by design, so it cannot speak to search, and the gap turns out to be wider when someone is actually hunting for something.
- The honest caveat is that this is 24 people, choosing professionally, from a fixed chart rather than a live scrolling feed. Nobody has published eye tracking of YouTube's real homepage or feed. That is a genuine hole in the research and worth knowing before anyone quotes a thumbnail statistic at you with confidence.
People use your thumbnail to decide what to skip

This is the finding almost nobody quotes, and it is in the same study. The researchers recorded people talking through their choices and coded which part of the card drove each decision.
- The thumbnail drove 26 decisions, and 21 of them were people deciding to skip a video rather than watch it. Only 5 went the other way. The video length behaved the same way, driving 23 decisions of which 18 were people skipping. The parts of the card people mention most are the parts they use to narrow the list down.
- When people did choose a video, they mostly could not say what on the card made them choose it. Out of 51 decisions to watch something, 31 could not be traced to any part of the card. People explained clearly why they skipped a video, and went quiet when asked why they picked one.
- This reverses how thumbnail advice is normally written. Almost every guide treats the thumbnail as a sales pitch that has to persuade somebody. What was actually measured is quieter than that. Its main job is to not be the obvious one to skip, and a thumbnail that is confusing, too small to read, or promising more than the video delivers fails that immediately.
- It also explains why so many thumbnail tests come back with no result. If most of what you gain comes from not being skipped, then two decent thumbnails will usually score the same, which is exactly what YouTube says happens in its own testing tool.
YouTube grades your thumbnail on watch time, never on clicks

YouTube has a built-in thumbnail testing tool, and the way it decides a winner contradicts nearly everything written about thumbnails.
- YouTube says outright that it does not judge the test on clicks. Its help documentation reads: "To help your video get high quality engagement, we optimize tests for overall watch time over other metrics, like click-through-rate." The winning option is the one with the highest watch time.
- A creator running YouTube's own thumbnail test is never shown a click rate at all. The results come back as watch-time share. So the metric the entire advice industry tells you to chase is the one metric YouTube's testing product declines to report.
- Getting no winner is normal and YouTube says so. Results come back as a winner, as "performed the same", or as inconclusive, and YouTube's own wording is that it is normal not to receive a winner. If your two thumbnails were both fine, that is the expected outcome, not a failure.
- The practical limits are worth knowing before you plan around it. You can test up to three options, a test usually finishes inside two weeks, and YouTube recommends starting with older videos to limit the damage if a variant does badly. Shorts, premieres, scheduled live streams and made-for-kids videos cannot be tested at all.
- If one of your test images is below 1280 by 720, every option in the test gets downgraded. A single low-resolution variant drags the whole experiment down to a lower quality, so one careless upload corrupts the result you were trying to measure.
A falling click rate is usually a sign your video is winning

This is the most useful thing YouTube publishes about thumbnails and it is buried in a help page. YouTube gives its own worked example of a click rate collapsing while the video succeeds.
- YouTube's example has a video go from 10,000 times shown at a 9% click rate to 100,000 times shown at 3.5%, and calls that a success. Their words are that in the context of 90,000 extra times shown, "this is a sign of success that your content is expanding to a broader segment of viewers". Do the arithmetic and 900 views became 3,500 views. The click rate fell by 61% while the views more than tripled.
- The reason is that your first viewers are your biggest fans and everyone after them is a stranger. YouTube states that early click rates are inflated because a video is first shown to people who already know your channel, and that the rate settles as the video reaches people who have never heard of you. A high click rate can simply mean your video never left your own audience.
- YouTube's Creator Liaison gives the same answer to what a good click rate is. Rene Ritchie wrote that "the only real, honest answer to this is, better than your last video", and illustrated it with a video getting 20% from a hundred superfans and 2% once the recommendation system showed it to a million people. Same video, same thumbnail, a tenfold difference in the number.
- The only official benchmark is a range, not a target. YouTube's published line is that half of all channels and videos sit between 2% and 10%. That is a description of the middle half of the platform, with no sample size and no methodology attached, and it is not an average. Any article quoting "the average YouTube click rate is 4 to 5%" has invented it.
- YouTube also names the pattern that means your thumbnail is overclaiming. A high click rate with a low average view duration and fewer times shown than you expected is the shape it describes as clickbait, and it says such videos are less likely to be recommended. That is the whole trade in one sentence: clicks you cannot back up cost you distribution.
The algorithm cannot see your thumbnail at all

YouTube's engineers published how the recommendation system works, and the thumbnail appears in it exactly once.
- YouTube's own recommender paper names the thumbnail as click variation the model cannot explain. In their 2016 paper on the deep neural networks behind YouTube recommendations, Covington, Adams and Sargin write that a user "may watch a given video with high probability generally but is unlikely to click on the specific homepage impression due to the choice of thumbnail image". That is the only mention of the word in the paper.
- No image feature appears anywhere in either of their models. The system uses video identifiers, search terms, demographics, location, device and timing. It does not look at the picture. So the thumbnail is a real source of clicking behaviour that the ranking system is explicitly blind to.
- That explains why YouTube grades thumbnails on watch time instead. If the recommendation system cannot read your image, the only way to judge it is by what people did after they clicked. The testing tool and the recommender are consistent with each other once you notice the algorithm was never looking at the artwork.
- It also means the thumbnail is the one part of the machine you control outright. Everything else about how your video is distributed is decided by a model reading signals you cannot edit. The thumbnail is not in that model, which makes it a lever rather than a lottery ticket, but only for the click, never for the recommendation.
More text on a thumbnail means fewer people watch

Text is the single largest gap between what creators are told and what has been measured, and the measurement is not close.
- The largest thumbnail study ever run found that more text went with less watching. Loebbecke and colleagues analysed roughly 500,000 thumbnails across two international platforms in the International Journal of Information Management in 2024. Their summary is that several visual cues and more faces encourage consumption "while more text decreases consumption". Nothing else in this field comes close to that sample size.
- A separate study of 16,215 YouTube videos found the same split between picture and words. Strong emotion in the thumbnail image went with more views, while strong emotion in the text written on the thumbnail went with fewer. Two independent studies now say the picture and the overlaid words behave differently, which is a good reason to stop treating them as one design decision.
- Punctuation, capitals and emoji are the one kind of textual decoration with a measured benefit. The same 16,215-video study found exclamation marks, all-capitals and emoji went with more clicking. Note this is typographic decoration, not more words, and the same paper found more emotive wording lowered views.
- For scale, 38% of thumbnails in that sample carried text at all and 51% carried a human face. So text on a thumbnail is common but not universal, which is worth remembering when someone tells you it is mandatory.
Nobody ever tested the three to five words rule

The most repeated rule in thumbnail design has no origin. I went looking for the study and there isn't one.
- The recommended word count is different everywhere you look, which is how you know none of it was measured. The numbers in circulation are 0 to 3, 1 to 3, 3 to 4, 3 to 5, 4 to 5 and "under 6", each stated with total confidence and none with a citation. A real measured optimum does not drift like that.
- YouTube has never published a word count. Its official thumbnail and title guidance says only that if you add text, use a font that is easy to read. That is the entire published guidance on the subject.
- The nearest thing to a source measured Facebook ads, not YouTube. A well-known strategist has claimed an 80% higher click rate from text versus no text, but the testing was run on Facebook advertising, with no sample size, no date and no method published, and his own conclusion was that it depends on the content rather than the word count. He never says three to five words.
- The number people cite as proof is a description, not a test. A 2026 study of 500 breakout videos reported that among thumbnails which already used text, the median was five words. That describes what popular creators happened to do. It has no comparison group, so it cannot tell you whether five words beats two or none.
- The rule that survives is legibility, not a word count. A thumbnail is often rendered around 168 pixels wide in a phone feed, with a duration badge over one corner and sometimes a red progress bar across the bottom. YouTube publishes no guidance about those overlays, and every "safe zone" percentage in circulation is a guess that contradicts the other guesses.
A face only helps where the other thumbnails have none

"Put a face on it" is the most confident advice in the subject. The evidence says it depends entirely on what the thumbnails beside yours are doing.
- Across more than 300,000 popular videos, thumbnails with and without faces performed about the same. That analysis covered 52,000 channels and 62.6 billion views, and its own summary was that faces amplify existing interest rather than create it. Faces hurt in gaming and television and helped in lifestyle, finance and beauty, and multiple faces beat a single face.
- The mechanism shows up cleanly in a marketing study of social media images. Li and Xie, in the Journal of Marketing Research in 2020, found adding an image lifted engagement on Twitter enormously, by 119% for air travel and 213% for cars, and did nothing at all on Instagram. Their explanation was that human images were already common on Instagram and so "may not be unique enough" to add anything.
- That inverts the statistic everyone quotes in favour of faces. The famous figure is that 69% of breakout thumbnails contain a face. It is a count of how common faces are, with no comparison group, published by a company selling thumbnail tools. Read properly it argues faces should help less, not more, because two thirds of your competitors already have one.
- Faces do reliably grab the eye, which is a different claim from getting the click. Eye tracking finds faces are looked at far more than matched control areas and are hard to avoid even when looking at them costs you something, and that about 70% of first looks land on a face despite faces filling a twentieth of the picture. Being looked at and being chosen are not the same thing.
- The exaggerated shocked expression has never been measured. A study of 500 breakout videos found only about one in twenty used one. The various "an expressive face lifts clicks 47%" figures trace to self-published marketing with no method, no venue and no sample.
Being vague only works when your rivals are being specific

The curiosity gap is the one piece of thumbnail folklore with a genuinely excellent study behind it, and the answer is not the one people expect.
- Being more specific raised clicks 5.5% against vague competition and cut them 9.9% against specific competition. A registered study in Scientific Reports analysed 8,977 headline experiments run on a site reaching about 50 million people a month. The effect flips sign depending purely on what the other options looked like.
- Only 8.7% of headlines gained from becoming more concrete, while 50.9% lost clicks. That asymmetry is the practical finding. Most of the time, adding specifics costs you, and the exception is when everything around you is being coy.
- So a curiosity gap is a positioning decision, not a technique. Being the one mysterious thumbnail among specific ones works. Being one more mysterious thumbnail among mysterious ones loses. That is why copying the style of whatever is winning in your subject tends to stop working the moment enough people copy it.
- The honest caveat is that this measured headlines, not thumbnails. It ran on a news site between 2013 and 2015, not on YouTube, and nobody has run the equivalent experiment on thumbnails. The mechanism is about competitive context, which transfers well, but the exact percentages should not be quoted as YouTube numbers.
Moderate detail beats a plain or a cluttered thumbnail

Complexity is the one design variable with proper causal evidence behind it, including randomised experiments rather than just observation.
- Thumbnail complexity follows an upside-down U, so both empty and cluttered lose to the middle. Fang and colleagues, in the Journal of the Academy of Marketing Science in 2026, studied 22,958 thumbnails and backed it with two randomised experiments. Very plain and very busy both underperformed a moderate amount of visual detail.
- YouTube's own guidance says the same thing in plainer words. Their thumbnail advice is that dynamic use of colour and composition can help catch the eye "but too much can overwhelm it". That is the same curve described without the maths.
- Cramming in more detail measurably hurts attention to your brand. Separate research across 249 advertisements found dense visual busyness reduced brand attention and made people like the advert less, while deliberate structure helped both. Busy is not the same as bold.
- The practical version is to add one strong element and then stop. The evidence does not support an empty minimal frame and it does not support four competing focal points. It supports a single clear idea with enough going on to look considered.
Arrows, circles and colour rules have no measurement at all

This is the shortest section in the article because there is genuinely nothing to report, and saying so is more useful than repeating invented percentages.
- Nobody has ever measured what an arrow, a red circle or a border does to clicks. Not a platform, not a vendor, not an academic. No study isolates them as a variable and reports a result with a real sample. Every "arrows lift clicks up to 25%" figure cites unnamed eye-tracking studies that are never named or linked.
- The two most quoted colour studies in the subject do not exist. A "TubeBuddy Design Study 2023" is cited for the claim that 85% of viral videos use highly saturated colour, and it does not appear anywhere on TubeBuddy's own site. A "Vidooly 2023 study" is cited for contrasting colours lifting clicks 30%, and no such report exists.
- The red-thumbnail numbers contradict each other, which is the signature of invention. The same claim circulates as 23%, 30%, 32% and 40%. A real figure degrades slowly as it is repeated. A fabricated one gets reinvented each time somebody needs it.
- The "consistency lifts clicks up to 38%" claim was traced and there is nothing at the end of it. The trail stops at the words "some studies" in blog posts, none of which predate 2025, and its attribution to a Wistia study is false, because Wistia's actual research measured something else entirely. The same 38% is reused across the same content ecosystem for faces, for neon accents and for consistency, which is how you know it is a stock number rather than a measurement.
- What is genuinely measured about colour is unglamorous and secondary. Where people look is predicted far better by what a thing means than by how much it stands out, and the gap is roughly 19% against 4%. Colour is worth something, but it is the smallest of the levers on this page.
Every thumbnail rule, graded by whether anyone measured it

This is the table I wanted and could not find anywhere. Green means real measurement with a meaningful effect. Amber means measured but small, mixed or indirect. Red means the advice is everywhere and the measurement is nowhere.
| Thumbnail rule | Evidence behind it | What was actually found | What to do about it |
|---|---|---|---|
| Make it readable at a glance | Eye-tracked, real YouTube layout | A fifth of a second per glance, glanced at four times as often | Do this first, nothing else survives if it fails |
| Differ from neighbouring thumbnails | 8,977 headline tests, pre-registered | Up 5.5% or down 9.9% depending on the competition | Look at the feed first, then decide your angle |
| Use less text, not more | 500,000 thumbnails, peer reviewed | More text went with less watching | Cut words, and stop chasing a word count |
| Moderate detail, neither bare nor busy | 22,958 thumbnails plus two randomised tests | An upside-down U, the middle wins | One strong idea, then stop adding |
| Put a face on it | 300,000 videos, and it depends | No overall difference, helps only where faces are rare | Check your competitors before assuming it helps |
| Use punctuation, capitals or emoji | 16,215 videos, peer reviewed | Went with more clicking, unlike emotive wording | Cheap to try, and it is decoration not words |
| Use a curiosity gap | Measured on headlines, never on thumbnails | Works only against specific rivals | A positioning choice, not a technique |
| Follow the rule of thirds | Tested once, and it failed | Good thumbnails put the subject anywhere | Ignore it, despite YouTube recommending it |
| Keep a consistent style across the channel | No measurement found | The famous 38% figure was invented | Do it for your own sanity, not for clicks |
| Add arrows, circles or borders | Never measured by anybody | Nothing exists, in either direction | Pure convention, so please yourself |
| Pick a specific colour, usually red | The two cited studies do not exist | Claims contradict at 23%, 30%, 32% and 40% | Stop testing colours, it is the smallest lever |
| Use an exaggerated shocked face | No measurement found | Only 1 in 20 breakout thumbnails even used one | Not the rule it is presented as |
- Four rules are worth acting on and every one of them is free. Make it readable at a glance, differ from your neighbours, cut the words, and land in the middle on detail. None of those needs a designer.
- YouTube officially recommends the rule of thirds and the only study that tested it found the opposite. Song and colleagues, analysing 1,118 videos, reported that good thumbnails contain a clear subject "in almost any position of the image" and so do not follow the rule of thirds. Worth knowing that the platform's own advice is unsupported by the one measurement of it.
- Read that study carefully before quoting it, because it never measured clicks. Its measure of a good thumbnail was which frame a professional editor picked. The authors say plainly in their conclusion that measuring click-through rates was still future work. It is routinely cited as evidence about clicking and it is not.
Nobody has tested a thumbnail against dark mode

Here is a hole big enough to drive a lorry through, and it is the one thing on this page you could genuinely settle yourself.
- Roughly half your viewers see your thumbnail against near-black rather than white, and nobody has ever tested that. YouTube's dark theme puts the surrounding page at nearly pure black, and its light theme puts it at white. No study, no YouTube statement and no documented creator experiment compares the same thumbnail across the two.
- Every dark mode statistic in circulation is invented. "82% of YouTube mobile users use dark mode", "dark thumbnails get a 15% visibility advantage", "dark backgrounds outperform light by 31% in gaming". None of these has a named study, a sample size or a method, and they appear only on content-marketing sites.
- You can settle it for your own channel in about two weeks with YouTube's own tool. Make one variant with a light background and one with a dark background, run the built-in test, and you will know more about your audience than any published research does. That is a rare position to be in.
- Bear in mind the tool will judge it on watch time, not clicks. So what you learn is which version brings people who stay, which is the more useful question anyway.
A banner fights to be seen, a thumbnail fights to be picked

I wrote a companion piece on what makes a scroll stopping banner, and its central finding looks like it contradicts this one. It concluded that trying to stand out is exactly what gets a banner skipped. On a YouTube feed, being different from your neighbours is one of the few things that works. Both are true, and the reason is worth understanding because it tells you which advice transfers between the two and which does not.
| What differs | A display banner ad | A YouTube thumbnail |
|---|---|---|
| Share of the time people looked | 5.5% while reading a page | 35.4% while choosing a video |
| What the viewer is doing | Trying to read something else | Actively choosing between options |
| Who it competes with | The page content itself | About twenty near-identical rivals |
| What standing out achieves | Marks you as an advert to skip | Makes you findable in the queue |
| Roughly how often it is clicked | Around 0.05% to 0.21% | Between 2% and 10% for half of channels |
| What the reader does with it | Filters it out before reading | Uses it to skip, four times out of five |
- The reconciliation is inside the very study the banner article was built on. Burke and colleagues also placed banners inside the area people were searching rather than at the top of the page. Those banners got looked at 58 times, against only 10 times for the ones at the top. The same rectangle, six times the attention, purely from sitting inside the region the person was already searching. Standing out never got the banner skipped. Being filed as outside the search was what did it.
- Standing out has no value on its own, only against competitors, and there is a clean experiment proving it. In the classic 1980 study of visual search, when only one item was on screen the eye-catching target was if anything slower to find, at 422 milliseconds against 426 and 446 for ordinary targets. The entire benefit of being distinctive is the benefit of being findable among rivals. A banner has no rivals to beat. A thumbnail has a screenful.
- The same banners got twice the attention when the task changed from reading to choosing. A study of 100 people found identical banners took 5.5% of viewing time when people were told to read the article and 11.9% when they were told to decide what to click next. A YouTube feed is permanently in that second state, which is why banner blindness barely applies to it.
- Being classified as content rather than advertising is what matters, not the shape or the colour. When researchers hid the information people needed inside a page's advert regions, success fell from 82% in the content area to 52.9% in the top advert strip and 36.8% in the side column. Almost everybody knew the side column held adverts; only half realised the top strip did, and it performed accordingly.
- But the returns on simply being louder collapse when every neighbour is loud too. Attention research finds a distinctive item stands out most when its surroundings are uniform, and a YouTube grid is the opposite of uniform. One model quantifies it: sharing a feature with half the things around you cuts its pulling power to a quarter. That is precisely what happens when every thumbnail in your subject adopts the same bold colour and shocked face.
- So the advice that transfers is "be legible and relevant faster than your neighbours", not "be louder than them". Where you sit still beats what you draw, in both mediums. The difference is that a banner is trying to breach a filter, while a thumbnail is trying to win a place in a queue, and only the second one rewards being different.
Conclusion: what actually makes a YouTube thumbnail work
- Four things are worth doing, and all four are free.
- Build it around one subject and a few large words, since one look lasts a fifth of a second.
- Make it different from the thumbnails beside it, not louder than them.
- Cut the words, and ignore every recommended word count.
- Aim for moderate detail, since both bare and cluttered lose to the middle.
- Your thumbnail's job is to not be the one people skip, rather than to persuade anybody. When people talked through their choices, the thumbnail drove 21 decisions to skip a video against 5 to watch one, and when they did pick something they mostly could not say what on the card made them do it. That is why a thumbnail somebody can take in at a glance beats a clever one, and why two competent thumbnails so often test the same. It also explains the failure mode: a thumbnail that overclaims does not just fail to win the click, it costs you distribution, because YouTube says videos with a high click rate and low watch time get recommended less.
- Stop optimising the click rate, because YouTube already refuses to. Its own thumbnail test never reports one and picks winners on watch time. Its own help pages teach that a falling click rate usually means the video is reaching past your existing audience, with a worked example turning 900 views into 3,500 while the rate drops by 61%. And its own engineers wrote that the recommendation system cannot see your thumbnail at all. Judge a thumbnail on whether the people it brought stayed.