



What Is Going On With Chess.com Ratings? How Much Has a Rating’s Meaning Changed Over Time?
Full article link: https://docs.google.com/document/d/e/2PACX-1vQyNolyIcUHBJaChN5-uMKZGxx2R1DmrsXGq6XR0Ci2nklabNQI1UqADoGvL9Eq0ihKS230NgiGoKz9/pub
General Key Findings:
- The rating distribution for Chess.com has shifted significantly downward over time, the distribution is now left-censored.
- The player base is composed of many more beginner-level players now than historically, justifying a shift in overall rating distribution shape to appear more right-skewed.
- Median ACPL is a good descriptor for overall playing strength and trends with overall rating regardless of time control or time period. Stronger players have higher ratings than weaker ones across the entire rating ladder.
- Median ACPL is converging towards some medium-value over time as stronger play is found at lower ratings, and weaker play is found at higher ratings as time progresses.
- Groups of players with the same median ACPL also have very similar CPL distributions in general.
- Lichess has much smoother move-quality distributions when compared to Chess.com, as Chess.com’s are very noisy and even have an observed trend reversal in the 3+0 time control.
- The outcomes of games on Chess.com, given both player’s ratings, are now much less predictable than historically.
- “Elo Hell” may exist not in the form of strong players being “stuck” at low ratings, but more so in that the relationship between player rating and playing strength has diminished significantly over time, especially for blitz in general and <1000-rated rapid.
- The distortion of the player strength to player rating relationship appears to begin at the 100-rating floor and move up the rating ladder over time, distorting the rating segments above it.
My theory:
I think, but don’t know for sure, that the surge in new players around COVID, also accompanied by their change to allow players to self-select their starting rating, either 400, 800, 1200, or 1600 must have led to many players that are much stronger than a 400 is “supposed” to be entering the pool at 400. This pushes down the bottom of the player pool significantly to the very bottom, 100.
Once players of varying beginner-level strength start being pushed down to 100, that means that “100” will cover a varied skill set. Which means that 150’s, being paired with 100’s will also see distortion in their ratings as a result. Then 200’s being paired with 150’s, and so on. Slowly over time the problem compounds and moves up the rating ladder as we see in the Glicko outcome error charts.