Reasoning Like Humans
Abstract
While recent reasoning models have achieved increasingly strong task-solving performance on ARC benchmarks, final correctness alone does not show that they reason in a human-like way. In our preliminary ConceptARC evaluations, current high-performing models remain much less reliable at intermediate concept-level reasoning. This gap motivates a process-level view of ARC evaluation: models should be assessed not only by whether they produce the correct output grid, but also by how they arrive at the answer. We formalize human-like ARC reasoning as a staged process of representation formation, problem-space selection, and heuristic search. Based on this view, we propose a Cognitive Reasoning Model composed of an Inductive Encoder, a Decision Decoder, and a Heuristic Searcher. By introducing a task-level concept prior before rule search, the model offers a concrete way to study structured reasoning in ARC beyond final-answer accuracy.