Large language models still struggle with many classic puzzle-style intelligence tests.
The magazine assembled sample problems to show where models misfire, highlighting spatial rotation tasks, visual riddles and trick questions that routinely trip them up spatial reasoning.
Benchmarks tell a mixed story: a Columbia team once found models solved only 18% of Connections puzzles, yet by early 2025 some systems could solve them near perfectly, illustrating rapid gain alongside persistent blind spots.
Researchers cited in the piece point to memorization and training-data overlap as failure modes, noting a 2024 Knights and Knaves study where small variations caused high-performing models to stumble.
The article also flags abstract-visual benchmarks such as ARC-AGI, where models improve when grids are provided as a string of numbers, suggesting they exploit encodings rather than human-like visual generalization.
Puzzles have been an evaluative tool since Arthur Samuel’s 1959 checkers work and Watson’s Jeopardy! run, and the field now layers game-based leaderboards and tougher robustness tests on top of old benchmarks as evaluators seek measures less prone to data contamination.
These examples don't change model capabilities overnight, but they sharpen where future evaluation and model design must focus.