Auditable Step Verification for Vision-Language Reasoning
Shervin Ghasemlou
Abstract
Outcome-only reinforcement learning from verifiable rewards cannot tell whether a vision-language model failed because it misread the image or because it reasoned incorrectly after reading it. We present $\textbf{VisualSPARK}$, a generative process reward model for multimodal math reasoning. VisualSPARK learns from Monte-Carlo-labeled reasoning steps and produces an auditable verification rationale ending in a $\texttt{[[correct]]}$ or $\texttt{[[incorrect]]}$ verdict. Across MathVista, MathVerse, MathVision, and MMMU-Pro, we evaluate on $77{,}680$ cached completions. The verifier reaches $81.3\%$ macro-F1 with $100\%$ parseable held-out verdicts. On $7{,}980$ free-form problems, PRM-weighted majority adds $+5.26$ aggregate points over pass@$1$---about $420$ expected additional correct selections---while preserving $98.6\%$ of majority-vote lift; pass@$8$ reveals $+25.57$ points of recoverable headroom. Diagnostics rule out length as a within-problem shortcut, quantify contamination and extraction sensitivity, and show practical batched serving. We release the full code, configs, result logs, datasheets, model cards, and measured compute accounting.
Chat is not available.
Successful Page Load