Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning
Noor Islam S. Mohammad ⋅ Mahmudul Hasan
Abstract
Large language models (LLMs) deployed for logical reasoning in knowledge-intensive domains frequently exhibit a subtle but critical failure: \emph{coherence can be vacuously achieved through systematic abstention}. A model that withholds commitment to either entailment or refutation trivially satisfies negation consistency while providing zero reasoning utility. We introduce \textbf{Coherence Under Commitment (CUC)}, a dual-query evaluation paradigm that jointly measures logical consistency and epistemic decisiveness. Our framework contributes three innovations: (1) a \textbf{commitment score} $c(\varphi) = p(\varphi) + p(\lnot\varphi)$ that quantifies probability mass allocated to decisive outcomes; (2) a \textbf{deterministic black-box elicitation protocol} via normalized YES/NO log-probabilities, eliminating sampling variance across runs; and (3) a \textbf{3-way decision framework} (\textsc{True}/\textsc{False}/\textsc{Uncertain}) that operationalizes the coherence-commitment trade-off into actionable metrics. Experiments on four open-weight LLMs (1B--3B parameters) across 204 \textsc{FOLIO} validation examples expose a sharp frontier: Qwen2.5-3B achieves near-zero contradiction ($\mathbb{E}[v_{\mathrm{neg}}]{=}0.025$) but only $7.4\%$ coverage, while TinyLlama-1.1B reaches $79.4\%$ coverage at the cost of negation violations on every example. Standard coherence-only evaluation would rank the systematically abstaining model first—CUC exposes this as vacuous. We argue that reliable reasoning evaluation requires reporting both coherence \emph{and} non-vacuous commitment and release a toolkit for standardized assessment.
Chat is not available.
Successful Page Load