What Lichess Analysis Misses
Lichess marked eight of my moves in a blitz game last week. Move 20 was not one of them, and it was a mistake, one of three in that game the report never mentions. The server had the position at +1.2 before 20.Qd1 and +0.9 after, too small a drop to mark. Stockfish given twenty times as long has it at +2.0 before and +0.7 after.
The same game evaluated twice. Lichess’s line is what its graph shows; the dots are the moves it marked, on both sides. The rings are the three the deeper search calls mistakes and it did not: 20, 24 and 26.
I wanted to know how often that happens, so I re-analysed 300 random games at twenty times the server’s budget, and 800 games of nine titled players at four times, and compared every label. One in ten changes. The inaccuracy label is the one to distrust. And every twentieth game has a mistake with no mark at all, which is where the puzzles come from.
How the labels are made
The server pass (fishnet, Stockfish on donated CPUs) evaluates each position in the game once, at 1,000,000 nodes a move, about a second of one core. Each label is the difference between two of those numbers, the position before your move and the position after, converted to winning chances so that a half-pawn slip at 0.00 counts for more than one at +5. A drop of 0.1 in winning chances is an inaccuracy, 0.2 a mistake, 0.3 a blunder.
So the label sees only what a million nodes see, on both ends. If the search never finds the refutation, the “before” number never knew what you had, and your move loses nothing against it. My 20.Qd1 was that case: the better square was one file over, and the punishment starts three moves later than a one-second search looks.
What changes at twenty million nodes
On 300 random rated games (1600 to 2400, from the August 2026 database), re-evaluated at 20,000,000 nodes with the same formula:
- One move in ten changes label. The changes go both ways in equal numbers, so the totals in the game summary barely move while the moves under them do.
- The inaccuracy is the unstable label. Of 1,805 Lichess inaccuracies, 57% are still inaccuracies at depth, 28% are nothing, and 15% are mistakes or blunders. Blunders hold 84% of the time.
- One game in twenty has a mistake with no mark at all, and one in eight has a marked mistake that is nothing at depth. Both cluster between half a pawn and three pawns, the range where games are decided.
The same pass over about a hundred games each of nine titled players’ own moves: a quarter of the mistakes Lichess pins on Magnus Carlsen are not mistakes at depth, and for everyone else it runs the other way, with one real mistake in five carrying a softer label than it deserves.
| player | games | Lichess mistakes+blunders | under-labelled | over-labelled |
|---|---|---|---|---|
| Magnus Carlsen | 104 | 195 | 18 (11%) | 50 (26%) |
| Alireza Firouzja | 104 | 215 | 45 (20%) | 35 (16%) |
| Eric Rosen | 102 | 225 | 45 (20%) | 43 (19%) |
| Jerry (ChessNetwork) | 104 | 287 | 45 (16%) | 46 (16%) |
| Oleksandr Bortnyk | 102 | 184 | 51 (25%) | 30 (16%) |
| Sergei Zhigalko | 103 | 218 | 53 (22%) | 29 (13%) |
| Dmitry Andreikin | 104 | 165 | 33 (19%) | 26 (16%) |
| Ediz Gürel | 61 | 117 | 37 (27%) | 15 (13%) |
| Anish Giri | 20 | 12 | 3 | 1 |
Five misses the analysis approved
Each position is from one of those games. The player made the move shown, the analysis gave it no mark, and a 20-million-node search calls it a mistake or worse against a best move that is unique. The board opens on the position before the move, so you can look for the better move first; step forward and you get the game move with the mark it should have had, and the refutation is the sideline in the study, one click under each board.
Zhigalko, Black to move. White has just played Be5, hitting the queen on c7, and Black did the natural thing: 20…Qd7, out of the attack. The analysis had it −1.38 before and −1.30 after, nothing lost. But the queen did not need saving. 20…Rxd2! gives up the exchange, and after 21.Rxd2 the pawn move 21…c3! forks the queen on b2 and the rook on d2 at once. White gets the queen back (22.Bxc7 cxb2 23.Rxb2 Rxc7) and is left with a rook against Black’s two bishops and knight: −3.2 at depth, against −1.3 for the retreat. The refutation starts with a move that loses material, which is why a one-second search never sees it.
Bortnyk, Black to move. 20…b5 attacks the bishop on c4, and the analysis moved from −0.63 to −0.41. The bishop on f5 could simply have taken the knight: 20…Bxe4! 21.dxe4, and now the other knight lands on f4, with …e5 to follow to push the bishop off d4. Black wins material inside five moves, −2.7 at depth; the second-best move, 20…Nf4 at once, is only −0.6.
Bortnyk, Black to move, a queen ending. Queen and knight against queen and rook, and White’s a-pawn is the only thing White has going. Black played 54…Qf6; the analysis said −1.58 to −1.03. At depth the position after Qf6 is a dead draw. 54…Qa7+! 55.Kg3 Qxa4 picks up that pawn with check, and the ending is −2.1 and winning.
Carlsen, Black to move. The white queen sits on g4 and the knight on e4. Carlsen played 15…Ne5, hitting the queen, and the analysis went from −0.94 to −0.43. The pawn move 15…f5! hits both: after 16.Qh3 fxe4 17.Bxe4 Black has won a knight for a pawn and stands −1.9, with the second-best move (15…g6) a pawn and a half behind. A shallow search reads …f5 as loosening the king and the knight retreat as safe, and both readings are wrong.
Gürel, White to move. A rook ending with a knight each; Black’s knight on a3 is loose, and White’s d-pawn is on d5. White played 31.Nd4, +1.42 to +1.06 by the analysis. 31.d6! is the move: the rook has to stop the pawn (31…Rxd6), and then 32.Nxa3 takes the knight for nothing. +3.9 at depth, against +0.9 after the game move.
All 32 are in the study, set as puzzles: the position first, the game move marked, the refutation as a sideline with the engine’s numbers, and every later move of the game that the 20M search judges carries its own mark and the engine’s line beside it. Forcing solutions come first, since a long tactic past the shallow horizon is something a person can find; the quiet ones are after.
Reading your own report
“Zero mistakes” means zero visible at a million nodes. The blunders are real, the inaccuracies are half real, and the moves that let an opponent back into a game you were winning by a pawn or two are the ones most likely to carry no mark. If a game matters, put your own engine on it and let it sit past depth 25 on the positions between +0.5 and +3.
How this was measured
The random sample is every tenth analysed game in the August 2026 Lichess database with both players between 1600 and 2400: 300 games, 20,227 moves with server evaluations. Each position went through Stockfish 18, the major version fishnet runs, at 20,000,000 nodes on one thread, and lila’s own formula (Advice.scala: winning-chance deltas of 0.1, 0.2, 0.3) produced a second set of labels. That formula, run on Lichess’s own evaluations, reproduces Lichess’s own labels 99.07% of the time, so both sides of the comparison use it. The 842 moves with a forced mate at either end are excluded; lila judges those by separate rules. The node budgets are in lila’s Work.Origin: 1,000,000 for a requested analysis, 5,000,000 for official broadcasts, 300,000 and 100,000 for the automatic passes.
| Lichess label | none at 20M | inaccuracy | mistake | blunder |
|---|---|---|---|---|
| none | 15,281 | 496 | 13 | 2 |
| inaccuracy | 504 | 1,033 | 233 | 35 |
| mistake | 27 | 190 | 367 | 151 |
| blunder | 12 | 18 | 134 | 889 |
The nine players’ games are their most recent hundred analysed rated blitz and rapid games from the public API (the active seven in 2026; Carlsen and Firouzja have no analysed Lichess games this year, so theirs run back to 2021 and 2023), at a cheaper 4,000,000-node pass. On the random sample, 4M finds a little over half of what 20M finds, so the player rates are floors. A puzzle survives when the move played loses at least 0.2 winning chances at 20M against a best move that beats the second line by 0.1 or more, with three lines searched, and then again when a second, independent 20M search of the position still calls the game move a mistake: 35 passed the first gate, 32 the second. That is the size of the noise in a 20M search, and it is why the deeper pass is treated as twenty times closer to the truth rather than as the truth.
Resources:
- The study: all 32 puzzles with solutions and engine numbers
- lila
Advice.scala: the judgment thresholds - lila fishnet
Work.scala: the node budgets per analysis origin - Lichess open database: the monthly dumps