SIEVE: When Small Language Models Look in the Mirror Limits of Inference-Only Self-Verification across Model Families and Tasks
Doguk Kim
Abstract
Agentic pipelines increasingly delegate correctness checking to the generator model, replacing human oversight with intrinsic self-verification. We ask when this delegation is safe for small language models (SLMs). We present SIEVE, a controlled study of six instruction-tuned SLMs (Qwen3.5 and Gemma4, 2B--31B) on GSM8K and HumanEval under four conditions: no verification (A), naive self-check (B), role-separated verification (C), and a high-capability external verifier (D, Gemini~3 Flash Preview). Three findings emerge. (i) On GSM8K, naive self-check produces a non-positive $\Delta$ for every tested model, with the magnitude statistically distinguishable from zero only at the lower end of the size tier (e.g., Qwen3.5-4B at $-15.5$pp; smaller drops for Qwen3.5-9B and Gemma4-E4B have CIs that overlap the baseline). (ii) Nominal gains for Gemma on HumanEval under self-verification have bootstrap CIs that overlap the baseline and remain unconfirmed. (iii) The external verifier is not a universal upper bound: with the prompt fixed, iteration-1 false-rejection rates on Condition-A-correct GSM8K answers span $0.0\%$ to $14.4\%$ across targets, and the same verifier exhibits up to $89.5\%$ false-acceptance on incorrect Gemma answers. We frame these as reproducible failure triggers and argue that SLM self-verification reliability is not a simple function of scale.
Chat is not available.
Successful Page Load