He's rated 1150. He's just lost, and he's explaining to you why that rating means nothing. He played blitz when he's really "a classical player." He hasn't worked on his openings in six months. His opponent won "on the clock, not on the board." And honestly, against the 1600 guy at the club, he wins "one in two anyway."

You know this player. Maybe you've been him. The interesting question isn't whether he's lying: he isn't, he sincerely believes what he's saying. The question is how many Elo points separate what he thinks he's worth from what he's worth.

Since August 2025, we have the number. It's 89 points.

What Kruger and Dunning Actually Measured

Let's start with the study everyone cites and almost nobody has read.

In 1999, Justin Kruger and David Dunning, then at Cornell University, published a paper in the Journal of Personality and Social Psychology with a title that pulls no punches: "Unskilled and Unaware of It." Four experiments, all built on the same design.

Students take a test. Then they're asked to rate themselves: "in your opinion, compared to the other participants, what percentile are you in?" Then the estimate is compared to the actual result.

The domains tested are deliberately varied. The first study looks at sense of humor: 65 students rate about thirty jokes, and their ratings are compared to those of a panel of professional comedians. The second covers logical reasoning, with questions taken from an LSAT prep manual. The third covers English grammar.

The result holds every time. Participants in the bottom quartile, the ones whose actual scores put them around the 12th percentile, place themselves on average at the 62nd percentile. Fifty points of gap. And note this: they don't think they're brilliant. They think they're slightly above average. Which is already enormous when you're in the bottom 12%.

The fourth study is the most important one, and it's the one everybody forgets. Kruger and Dunning take 140 students. They train 70 of them for ten minutes in logical reasoning, and keep the other 70 busy with an unrelated task. Then they ask everyone to rate themselves again. The trained students revise their self-assessment downward. Not because anyone told them they were bad, but because they've just acquired enough to see what they were missing.

That's the paper's real thesis. It isn't "incompetent people are too stupid to know they're stupid." It's far more elegant: the skills that let you do something well are the same skills that let you judge whether you're doing it well. When they're missing, they're missing twice. Once for the performance, once for the assessment of the performance.

Put that way, in chess it becomes almost obvious. What lets you find the right move and what lets you recognize that your move was bad come from the same body of knowledge. The 1100 player doesn't just miss the winning move: he also misses the fact that he missed it.

The Curve You've Seen Everywhere Is a Fake

Before going further, something has to be cleared out of the way, because it poisons every discussion of the subject.

You've definitely run into the graph. A curve that shoots up toward a summit called "Mount Stupid," collapses into a "valley of despair," then climbs gently along a "slope of enlightenment" toward a "plateau of sustainability." Experience on the x-axis. Confidence on the y-axis.

That graph does not appear in the 1999 paper. Nor in any publication by Kruger and Dunning. It's an illustration born on the internet in the 2010s, retroactively attributed to the study by sheer force of sharing. The terms "Mount Stupid" and "valley of despair" are not their vocabulary.

And this isn't a cosmetic detail, because the fake graph tells a different story from the real data, on two counts.

First, the axes aren't the same. In Kruger and Dunning, the x-axis is the quartile of actual performance at a given moment, not progression over time. The y-axis is an estimated percentile, not some floating "confidence." The viral graph turns a photograph into a film, which is not the same claim at all.

Second, and this matters more: in the real data there is neither peak nor valley. Perceived performance rises steadily with actual performance. The weak overestimate themselves, the strong are roughly accurate, and between the two the curve climbs quietly. Nobody collapses. Nobody peaks at 1200 Elo in a delirium of confidence before crashing at 1600.

If you've read somewhere that the club player goes through a "valley of despair" around 1600 Elo, you've read a nice story, not a result. The idea is appealing, it resembles an experience many players think they recognize in their own progress, and that's exactly why it circulates so well.

What the research says is simpler, and more uncomfortable: overconfidence doesn't peak, it declines. It's maximal right at the bottom and fades out gradually.

2025: Chess as a Laboratory

Here's the problem with the whole Dunning-Kruger literature: it rests on American undergraduates taking a grammar test on a Tuesday afternoon. Those people have no objective feedback on their grammar skills. They've never received a reliable, public, continuously updated score. It's easy to overestimate yourself in the dark.

Hence the obvious criticism: would the effect survive if people had an exact, permanent, public measure of their competence?

As it happens, one domain offers exactly that. One, or close to it.

In August 2025, Patrick Heck, Daniel Benjamin, Daniel Simons and Christopher Chabris published a study in Psychological Science titled "Overconfidence Persists Despite Years of Accurate, Precise, Public, and Continuous Feedback: Two Studies of Tournament Chess Players." The title is already a summary.

The setup is simple. Two preregistered studies, one of them a replication. In total 3,388 rated players, aged 5 to 88, average age 45, from 22 countries, with an average of 18.8 years of tournament experience. These aren't beginners surveyed in a corridor: they're club and tournament players, federated, with nearly two decades of perspective on their own curve.

They're asked for their current rating. Then they're asked whether that rating faithfully reflects their level. Those who say no have to specify in which direction, and above all give the rating they should have if it were accurate.

That's a clever piece of method. The player isn't asked to place himself on some vague scale: he's asked to produce a number, in a unit he's been handling for years, and one that's directly comparable to an objective value already on record.

The Number

On average, players claim a level 89 Elo points above their observed rating.

Take a second to translate what that means, because Elo isn't an intuitive scale. An 89-point gap corresponds to an expected score of roughly 62.5%. In other words, these players are implicitly asserting that against an opponent with exactly their current rating, they should score five points out of eight. Five wins for three losses. Against their own reflection.

At the group level, that's arithmetically impossible. If everyone is worth 89 points more than their rating, then nobody is worth 89 points more than their rating.

The Dunning-Kruger Pattern

The gap isn't uniform. It's largest among the lowest-rated players and smallest among the highest-rated. Players at the top of the table are calibrated: what they think they're worth matches what they're worth.

That's exactly the structure described in 1999, but measured this time against an external yardstick rather than students' cross-assessments. And the pattern shows up in every sociodemographic subgroup examined. Not an age effect, not a country effect, not a gender effect.

The One-Year Verdict

This is the harshest part of the result, and the most useful.

The researchers went back to check the ratings a year later. Were the overconfident players right? Were they simply ahead of their rating, mid-climb, legitimately convinced the number would catch up with the level?

11.3%. Slightly more than one in nine did reach the rating they'd claimed. The others improved, many of them, but far less than they'd announced.

So the objection "I'm underrated, it'll show soon" is true one time in nine. Statistically, when you say it, you have eight chances out of nine of being wrong.

Why Chess Should Have Killed This Bias

You have to appreciate how strange this result is.

Competitive chess is the most hostile environment to overconfidence you could design. The Elo rating is objective, computed by a public formula. It's precise, to the point. It's public, viewable by anyone on the federation's website. It's continuous, updated after every tournament. And it's predictive: it announces in advance the expected score against a given opponent, and that prediction holds up.

No other area of life works like that. Your employer doesn't hand you a four-digit number refreshed every month. Your skill at driving, at cooking, at management is never measured with that kind of arithmetic brutality.

A player with 18.8 years of experience has taken that feedback thousands of times. And he still believes he's 89 points above it.

How is that possible? Three mechanisms, and they stack.

The Rating Gets Reinterpreted, Not Accepted

The feedback is received, but it isn't processed as a measurement. It's processed as an opinion open to debate.

The rating measures a score under given conditions. The player knows plenty of things the rating doesn't: that he was tired in round 5, that he plays better when he's prepared, that he was going through a rough month, that he lost two winning positions on the clock. Each of those pieces of information is true. And each supplies a legitimate reason to consider the number an underestimate.

The problem isn't the information, it's the asymmetry in how it's used. Those circumstances get mobilized to explain losses, almost never to put wins in perspective. Nobody says "I won that game but my opponent was ill, so my rating overrates me." That one-way handling has a name, and it's the subject of another article in this series, coming soon: rationalization.

Memory Doesn't Store a Representative Sample

Ask a club player to tell you about his games. He'll tell you about the win against the player rated 300 points above him. He won't tell you about the eleven losses to players at his own level.

This isn't bad faith. The memories that stand out are the ones that break the routine, and by construction, your best games break the routine. Your mental stock of examples is therefore skewed upward, structurally. When you assess yourself from your memories, you assess yourself from your ceiling, not your average.

The Elo rating, meanwhile, is an average. It can't possibly agree with your memory.

The Ceiling Is Invisible From Below

This is Kruger and Dunning's original mechanism, and it's the deepest one.

A 1200 player watches a game between 1800 players. He sees moves. He understands the moves, one by one. None of them looks out of reach: they're legal moves, often natural, sometimes obvious once played. He concludes, sincerely, that there's "not that much of a gap."

What he doesn't see is what wasn't played. He doesn't see the four plans discarded in twelve seconds because the 1800 player knows the resulting structures. He doesn't see the line calculated and then abandoned. He doesn't see the natural move that wasn't played because it weakens a square ten moves down the road.

The gap between 1200 and 1800 is made of things absent from the board. And you can't measure an absence you don't know how to identify. That is precisely the 1999 thesis: you need the skill to see what the skill makes possible.

The Counterattack: What If the Effect Doesn't Exist?

Now, honesty demands the part most Dunning-Kruger articles quietly skip. The effect is seriously contested.

Starting in 2016, Edward Nuhfer and colleagues published two papers showing that the 1999 analysis method produces the Dunning-Kruger pattern even on purely random data. In 2020, Gilles Gignac and Marcin Zajenkowski drove the point home in the journal Intelligence with an unambiguous title: the Dunning-Kruger effect is "(mostly) a statistical artefact."

The argument comes in two pieces.

The first is regression to the mean. When two variables aren't perfectly correlated, and self-assessment never is with performance, extreme values on one correspond mechanically to less extreme values on the other. The weakest on the test overestimate themselves, the strongest underestimate themselves. No psychological phenomenon required.

The second is the better-than-average effect. Everyone, at every level, tends to place themselves slightly above average. Combine that constant upward shift with regression to the mean, and you get the 1999 figure without needing to invoke metacognition at all. In 2022, the analyst Blair Fix popularized a more radical version of the criticism by showing that part of the pattern comes from autocorrelation: relating a variable to a quantity that already contains it.

Dunning replied in 2022, in substance: purely statistical explanations neglect converging results obtained through other protocols, and the fact that people assess themselves badly remains true whatever the mechanism.

Why the Chess Study Changes Things

Here's where the case gets interesting for a chess player.

A good part of the statistical criticism targets one specific point: the measure of competence and the measure of self-assessment come from the same test, administered at the same moment, with the same measurement error. That common origin is what creates the artifact.

The 2025 study doesn't work that way. Actual level isn't a score obtained that day: it's an Elo rating built over years and hundreds of games, measured independently, by a third party, before the question was ever asked. Self-assessment is collected separately. And validation doesn't come from a correlation but from a one-year longitudinal follow-up with a binary criterion: did the player reach the rating he claimed, yes or no?

That architecture doesn't make the study invulnerable, none is. But it dodges the heaviest objections aimed at the 1999 protocol. And it's no accident that chess is what made it possible: it's one of the rare human domains where an external, longitudinal, reliable measure of competence exists.

Put another way, while psychology was tearing itself apart over the validity of Dunning-Kruger, the cleanest answer was sitting on a chessboard.

What This Changes at Your Board

Let's get concrete. Five levers, from the simplest to the most demanding.

1. Take the 89-Point Test

Write down right now, before you read on, the rating you think you deserve. Note the gap with your actual rating.

Then translate it. A 50-point gap implies you should score roughly 57% against a player at your current rating. A 100-point gap, about 64%. A 200-point gap, about 76%.

Now open the history of your last fifty rated games against opponents at your level, plus or minus 50 points. Count. The real score is right there, it has no ego, and it settles the matter.

2. Separate "I Can" From "I Do"

Overconfidence often comes from confusing peak capability with average performance. You can find that mate in three. That's not the question. The question is: across a hundred comparable positions, in a real game, with the clock ticking and the fatigue of round 4, how often do you actually find it?

Elo measures your average. Your intuition measures your ceiling. The two will never meet.

3. Have Your Games Analyzed by Someone Stronger

The engine tells you that the move is bad. It doesn't tell you what you failed to see, or why you couldn't see it. And that's exactly the blind spot Kruger and Dunning described: the thing that's invisible from below.

It's also the lesson of their fourth study, the only one that offers a remedy. Ten minutes of training were enough to recalibrate self-assessments. Not ten minutes of lecturing: ten minutes of genuine learning. Clear-sightedness doesn't come from wanting to be clear-sighted, it comes from learning. A player 400 points above you who goes through one of your games will teach you more about your real level than six months of solo engine analysis. The full method will be covered in our upcoming guide on analyzing your own games.

4. Keep a Log of Your Reasons

After every loss, write down in one line the reason you give. Don't try to be objective, write whatever comes to you first.

Read it back after thirty games. If "no time," "tired" and "unlucky" account for most of your lines, you're holding written proof of the reinterpretation mechanism. One isolated reason explains one game. Thirty isolated reasons describe a level.

5. Accept the Translation Both Ways

This is the hardest point. The rating isn't an opinion about your worth, it's a summary of your results. It doesn't say what you're worth as a player in potential, or what you could be worth with work. It says what you scored.

There is a healthy way to live with that sentence. It consists of no longer reading it as a verdict. That will be the subject of our upcoming article on Elo and self-esteem, and it's probably the real issue behind overconfidence: you're not defending a number, you're defending the image you have of yourself.

The Final Twist: What If You're in the Other Camp?

One last point, because this article could give the impression that everybody overestimates themselves. That isn't what the study says.

The highest-rated players don't underestimate themselves: they're accurate. The pattern isn't symmetrical. There's no mass of modest strong players offsetting the mass of presumptuous weak ones. There's massive overestimation at the bottom that fades gradually as you climb.

Which means that if you're a solid player who constantly doubts his own legitimacy, you're not a Dunning-Kruger case. You're something else, something that touches on a sense of belonging rather than an assessment of skill, and it's called impostor syndrome.

The two phenomena look like two sides of the same coin. They aren't. One is an estimation error, the other is a relationship with yourself. You can know your level to the point and still not feel like you belong.

Key Takeaways

The Dunning-Kruger effect as it circulates on the internet, with its Mount Stupid and its valley of despair, doesn't exist. The graph is a fake and the curve it describes matches no data.

The effect as the research documents it is quieter and more uncomfortable: the weakest players overestimate themselves the most, and feedback isn't enough to correct them. In chess, where that feedback is accurate, precise, public and permanent, the average gap between claimed level and actual level reaches 89 Elo points after nearly twenty years of practice. A year later, fewer than 12% of overconfident players had caught up to the rating they'd assigned themselves.

The only remedy identified in the literature isn't modesty. It's learning. You don't become clear-sighted by deciding to be: you become clear-sighted by acquiring what was missing, the thing that made the blind spot invisible.

Which is, when you think about it, a very solid reason to keep working.

After reading: take the 89-point test. Write down the rating you think you deserve, calculate the score that gap implies, and hold it up against your last fifty games against opponents at your level. The result, whatever it is, is more instructive than this article.

This article opens a series devoted to psychology applied to chess. It will be followed by articles on cognitive dissonance, confirmation bias, defense mechanisms and rationalization. See also the overview of the 5 cognitive biases that make you blunder. See the full index of the series.

Sources

  • Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it : How difficulties in recognizing one's own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121-1134.
  • Heck, P. R., Benjamin, D. J., Simons, D. J., & Chabris, C. F. (2025). Overconfidence persists despite years of accurate, precise, public, and continuous feedback : Two studies of tournament chess players. Psychological Science, 36(9), 732-745.
  • Gignac, G. E., & Zajenkowski, M. (2020). The Dunning-Kruger effect is (mostly) a statistical artefact : Valid approaches to testing the hypothesis with individual differences data. Intelligence, 80, 101449.
  • Nuhfer, E., Cogan, C., Fleisher, S., Gaze, E., & Wirth, K. (2016). Random number simulations reveal how random noise affects the measurements and graphical portrayals of self-assessed competency. Numeracy, 9(1), article 4.
  • Krueger, J., & Mueller, R. A. (2002). Unskilled, unaware, or both? The better-than-average heuristic and statistical regression predict errors in estimates of own performance. Journal of Personality and Social Psychology, 82, 180-188.
  • Dunning, D. (2022). The Dunning-Kruger effect and its discontents. The Psychologist, British Psychological Society.
  • Jarry, J. (2020). The Dunning-Kruger effect is probably not real. McGill Office for Science and Society.