Article
AI News AI Models & Research

AI models stumble on classic puzzle tests

MIT Technology Review showcased a set of puzzles that illustrate how current large language models can excel on many benchmarks yet still fail at spatial, visual and trick-based problems where humans retain an edge.

by TechDefused Newsroom
The image features a close-up of a puzzle with a piece missing, illuminated by light coming through the gap. The surrounding puzzle pieces are in shadow, emphasizing the absence of the missing piece. — Credit: Photo by Edge2Edge Media / Unsplash cPhoto by Edge2Edge Media / Unsplash
Photo by Edge2Edge Media / Unsplash

Large language models still struggle with many classic puzzle-style intelligence tests.

The magazine assembled sample problems to show where models misfire, highlighting spatial rotation tasks, visual riddles and trick questions that routinely trip them up spatial reasoning.

Benchmarks tell a mixed story: a Columbia team once found models solved only 18% of Connections puzzles, yet by early 2025 some systems could solve them near perfectly, illustrating rapid gain alongside persistent blind spots.

Researchers cited in the piece point to memorization and training-data overlap as failure modes, noting a 2024 Knights and Knaves study where small variations caused high-performing models to stumble.

The article also flags abstract-visual benchmarks such as ARC-AGI, where models improve when grids are provided as a string of numbers, suggesting they exploit encodings rather than human-like visual generalization.

Puzzles have been an evaluative tool since Arthur Samuel’s 1959 checkers work and Watson’s Jeopardy! run, and the field now layers game-based leaderboards and tougher robustness tests on top of old benchmarks as evaluators seek measures less prone to data contamination.

These examples don't change model capabilities overnight, but they sharpen where future evaluation and model design must focus.

by TechDefused Newsroom