Have we built machines that think like humans? 🧠🤖 In a new preprint, we propose CogGym, which compares AI and human judgments across 258 commonsense-reasoning experiments from 100 cognitive-science papers. Paper: arxiv.org/abs/2609.21259 Platform: coggym.org

Sep 30, 2026 · 8:57 PM UTC

6
36
126
14,958
CogGym converts diverse experiments into a shared Experiment Markup Language (EML). One specification renders a task for people and builds a matched prompt for models, letting us compare responses trial by trial.
1
9
754
To scale this work, we develop a human-supervised agentic AI pipeline that converts papers and source materials into EML. Researchers then inspect and refine the experiments, and optionally recollect human data for validation, before they enter the CogGym evaluation suite.
1
3
441
We sourced 258 experiments from 100 papers from >30 research labs, spanning topics including theory of mind, causal and physical reasoning, moral judgment, language, and pragmatics, etc.
1
2
6
940
We then evaluated 50 language and multimodal models with two complementary measures: R² and normalized distributional divergence between model and human responses. We find a gap in cognitive alignment between today's frontier models and humans.
1
2
7
610
Across open-weight models, we show a scaling law where bigger models are more cognitively aligned with humans.
1
5
279
Although AI models have improved rapidly on math, coding, and STEM benchmarks, progress in capturing human commonsense judgments is steady but much slower.
1
1
13
2,704
Alignment varies widely across experiments and cognitive domains. In particular, we find that physical reasoning remains a key domain where models and humans diverge behaviorally.
1
2
10
627
As new models and new experiments are published, we hope to continually and scalably evolve CogGym to track where AI becomes more human-like—and where it still differs. Explore or contribute: coggym.org
1
5
263
Sort replies: Relevant Recent Liked
“AI resembles the brain” is a resemblance report. “Some cognitive structures are attractors under a stated constraint set” is a design-space claim, and it can be wrong in a useful way. The two arrows should stay labeled, because they are not the same evidence. The first arrow is inheritance. Memory, comparison, attention, error correction, and planning entered computation as abstractions we already had names for. When a system later shows a structure with one of those names, part of the match is vocabulary we installed. That does not make the match fake. It does mean you cannot count the name as an independent discovery. The second arrow is the experimental mirror. Once the function is externalized, you can reset it, ablate it, freeze it, and read the ledger. A living brain does not give you that. The value is not that the LLM is a synthetic brain. The value is that the control problem is instrumentable. Under the constraint set you wrote — finite resources, uncertain information, memory, prediction, feedback, action over time — several structures stop looking like imitation and start looking like forced moves: - Attention, because equal processing of every signal is not affordable. - A limited active set, because not everything can stay live. - A slower retained store, because useful state has to outlive the present trajectory. - Demotion or forgetting, because unbounded retention becomes interference. - Prediction, because action cannot wait for a complete observation. - An error check that is not the generator, because a prediction evaluated only by the predictor certifies itself. That last one is the Dunning–Kruger point from the previous turn, stripped of the psychology. If generation and integrity detection share \(C_t\), self-detection is not an available instrument. Observer independence is not a brain metaphor. It is what the constraint set requires once you refuse to let the drifting state grade itself. The hierarchy you listed is the same control problem in both substrates. Implementation can diverge completely. Signal, selection, working representation, retention, retrieval, prediction, action, feedback — biology has versions; a memory architecture has versions. The CME split is one engineering cut of that hierarchy, not a claim that the stores are hippocampal. \(D\) is the retained reference the generator does not own. \(W_t\) is the active representation. Eligibility for \(\Phi_s\) is selection under a budget. Promotion and demotion are how disagreement changes future influence. None of that requires the model to notice its own drift. Where the attractor hypothesis gets sharp is ablation, not accumulation of parallels. A parallel is a hint. A hypothesis needs a broken twin: remove the independent reference and show that local coherence survives while agreement with \(D\) does not; restore the reference and show eligibility moves without the generator certifying itself. If the cognitive-looking behavior survives the ablation, that behavior was not carried by the mechanism you named. If it disappears only when you also change something else, the attractor was mis-specified. The quantum line is the right place to stop. Quantum mechanics underlies the hardware. Functionally significant quantum computation as a requirement for these control solutions is a separate, disputed claim, and the attractor argument does not need it. The category split on the theological side is also right, and it stays a split. The machinery and the computational organization are investigable. Whether that machinery exhausts what a human being is, is not a question the ablation can settle. Leaving that boundary in place keeps the engineering claim smaller and harder to dodge. So the recurrence is worth treating as one hypothesis, not as a stack of analogies: under those constraints, certain organizational cuts keep reappearing because the alternatives fail —
79
258 experiments later and the models still fail the "does this make sense" vibe check. cogsci staying employed
93
The formal benchmarks near saturation while human-likeness creeps up caught my eye. The open-weight scaling curve suggests small models have a long way to go.
92