Puzzles and games have occupied a central role in the evolution of artificial intelligence since its very inception. Just as humans enjoy testing their cognitive agility with crosswords, riddles, or logic puzzles, software developers and computer scientists have long relied on gaming gauntlets to measure how far machine learning models have advanced. The term "machine learning" itself was popularized in a landmark 1959 article written by IBM computer scientist Arthur Samuel, detailing an algorithm that successfully learned to play checkers. Decades later, complex strategy boards like chess and the ancient Chinese game of Go became legendary test beds for evaluating the upper limits of computational capability.
Judged purely on their raw puzzling skills, modern artificial intelligence models are improving at a remarkable and unprecedented pace. In late 2024, a team of researchers from Columbia University demonstrated that even the most sophisticated frontier models at the time could successfully solve only about 18% of the notoriously difficult New York Times Connections puzzles. Yet, by early 2025, newer iterations of these models had evolved to solve the exact same linguistic categorization challenges near perfectly on every single attempt, highlighting a staggering trajectory of capability growth.
However, puzzles serve a much deeper purpose than merely showcasing the inexorable advance of artificial intelligence. Pinpointing the exact areas where models succeed and fail—and contrasting them with domains where humans continue to consistently outperform machines—provides a uniquely revealing window into the underlying strengths and weaknesses of contemporary technology. Despite monumental leaps in computational power and training data volume, today’s models still fumble in predictable ways. Subtle alterations in classic riddles routinely trip them up, and visual puzzles remain a particularly stark weak spot.
Examining these vulnerabilities offers a fascinating opportunity to test human wits against problems that have routinely confounded algorithmic models. While some of these challenges prove just as tricky for people as they are for artificial intelligence, others are deceptively simple, prompting observers to question whether machine intelligence functions anything like human cognition at all. Each problem highlights distinct divergences between machine and human thought processes.
Spatial Reasoning
One domain where human cognition retains a massive and persistent advantage over silicon is spatial reasoning. For individuals who have ever undergone cognitive assessments or IQ testing, mental rotation problems are a familiar hurdle. These spatial puzzles require the test-taker to determine whether multiple distinct images represent the exact same physical object viewed from different angles.
Although contemporary multimodal language models possess the technical capability to ingest and analyze visual inputs, they still fail abysmally when confronted with these spatial orientation tests. Despite widespread industry discourse regarding how advanced "world models" can help artificial intelligence comprehend physical environments, large language models (LLMs) fundamentally struggle to mentally manipulate three-dimensional objects in the fluid manner native to architects, mechanical engineers, and everyday spatial thinkers.
Memory and Adaptability
Frontier large language models boast extraordinary memory capacities, having been exposed to an astronomical volume of factual data during their training phases. They can faithfully recite vast quantities of this information upon command. While this massive repository of memorized facts is an undeniable asset when competing against humans in general trivia contests, it can simultaneously act as a severe liability.
When a logical puzzle closely resembles a problem structure that a model encountered frequently during its training phase, the algorithm frequently races past critical, nuanced differences. Instead of performing genuine logical deduction, the model often defaults to regurgitating a memorized template or pattern.
This phenomenon was robustly documented in a 2024 study where researchers from Google and the University of Illinois Urbana-Champaign trained and tested various models on slight, deliberate variations of a classic puzzle genre known as Knights and Knaves. Within these traditional logic problems, certain characters always tell the absolute truth while others perpetually lie, requiring the solver to deduce the true identities of each participant based entirely on their statements.
A similar algorithmic pitfall appears to be at play in an evaluation framework called SimpleBench. These specific questions resemble more complicated, highly technical problems that models likely encountered extensively during their pre-training phases. While human solvers effortlessly spot the underlying tricks and conceptual pivots, even top-tier commercial models routinely trip over the deceptively simple phrasing.
Abstract and Visual Reasoning
Artificial intelligence does not merely stumble when processing three-dimensional visual dilemmas; two-dimensional abstractions can prove equally confounding. This limitation plays a major role in determining how successfully models navigate the most famous puzzle-based benchmark in the artificial intelligence community, known as ARC-AGI. These problems demand that the solver infer abstract, generalizable rules from a limited set of visual or structural examples.
Interestingly, research indicates that models perform noticeably better on ARC-AGI puzzles when the grid data is fed into them not as a traditional pixel image, but rather as a string of numbers explicitly encoding the color coordinates of each individual cell.
Furthermore, scientific studies suggest that even when models occasionally arrive at the correct answer for an ARC-AGI challenge, they frequently do so by relying on byzantine, highly specific, and ultimately non-generalizable computational rules. Humans, by contrast, naturally draw upon simple, elegant visual concepts and universal design intuition. Despite these inherent architectural disadvantages, artificial intelligence models have demonstrated notable performance gains on ARC-AGI over the past year, though a distinct subset of complex spatial-transformative puzzles continues to completely stump them.
Intuition
It is not merely artificial intelligence models that fall victim to predictable cognitive traps; human beings possess their own deeply ingrained cognitive foibles—many of which machines do not share. In fact, psychologists have designed specialized problem suites that essentially invert the SimpleBench phenomenon. In these particular categories of questions, human solvers frequently default to rapid, knee-jerk intuitive answers, whereas advanced models respond in a more calculated, deliberative fashion.
Some of these psychological test problems intentionally exploit systemic errors in how humans intuitively approach basic mathematical scaling. Others are phrased deliberately to suggest an obvious narrative answer that entirely falls apart if the underlying text is read with careful, literal scrutiny.
Increasing Complexity
In many practical scenarios, whether a large language model can successfully complete a complex puzzle is fundamentally a question of scale and resource allocation. A prominent study conducted by researchers at Apple revealed that contemporary LLMs can easily ace simplified versions of classic computational puzzles. These include the Tower of Hanoi problem—which requires moving a stack of graduated disks one at a time without ever placing a larger disk atop a smaller one—as well as classic river-crossing puzzles, where a mixed group of travelers must navigate across a body of water subject to strict behavioral rules.
However, this competency holds true only up to a strict threshold. As the scale increases—such as when the number of disks in the Tower of Hanoi or the number of passengers in a river-crossing scenario climbs to six or higher—the models begin to systematically falter and produce invalid sequences.
In a separate line of inquiry, researchers hailing from the University of Washington, Stanford University, and the Allen Institute for AI observed that large language models struggle with remarkably similar failure modes when confronting standard logic grid puzzles. These exercises require the solver to patiently deduce the precise attributes of a set of individuals by processing a lengthy, interconnected list of clues.
While the Apple research paper circulated widely across tech communities, independent commentators and computer scientists frequently debated whether these findings expose a uniquely catastrophic limitation in large language model reasoning architecture, or if they simply demonstrate that making errors under mounting complexity is an inherent byproduct of scaling computational systems.
As researchers continue to refine evaluation frameworks and push the boundaries of machine reasoning, the gap between human intuition and algorithmic computation remains narrow in some domains yet stubbornly wide in others. Whether artificial intelligence will eventually master every nuance of human puzzle-solving or continue to stumble over the simplest tricks remains one of the defining questions of contemporary computer science.