A Multi-Model Self-Evolving Framework for Zero-Data Document Understanding via Axiomatic Synthetic Refinement
Abstract
The scalability of Vision-Language Models (VLMs) for structured data understanding is fundamentally bottlenecked by the scarcity of high-quality, human-curated datasets. We present Omni-Zero-Doc, a self-evolving framework designed for zero-resource document intelligence. By orchestrating a tri-role agentic system—Proposer, Coder, and Solver—integrated with the Style-Isomorphic Evidence-Grounded Evolution (SIEGE) engine, we enable models to programmatically synthesize high-fidelity, structure-aware visual data. To ensure logical consistency, we employ Group Relative Policy Optimization (GRPO) coupled with a multi-stage verifiable reward mechanism and axiomatic filtering to eliminate hallucinations. Across benchmarks including DocVQA, MathVista, and MMMU, Omni-Zero-Doc achieves a +14.0\% ANLS gain, demonstrating that iterative self-play can sustain performance without external supervision.