Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification
Abstract
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. UniAR adapts a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, preserving both high-level semantics and low-level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel-bitwise-prediction to jointly predict spatially grouped, multi-level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion-based visual decoder operates on discrete visual tokens to decode high-fidelity images. Through large-scale pre-training, followed by supervised fine-tuning and reinforcement learning, UniAR achieves state-of-the-art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project homepage is at https://sharelab-sii.github.io/uniar-web.
Lay Summary
Current AI systems that can both understand and create images typically rely on two separate ways of representing visual information, one for comprehension and another for generation. This split prevents the model from directly interpreting images it has just created, limiting true unification of these abilities. We introduce UniAR, a system that uses a single, shared image representation for both seeing and creating. This allows the model to seamlessly understand and generate visual content within the same conversation, much like how humans use one visual system to both perceive and imagine. Our compact representation captures both fine details and high-level meaning, enabling efficient generation of high-quality images. Through large-scale training and learning from feedback on its own outputs, UniAR achieves state-of-the-art performance in image generation and editing, with particularly strong text rendering, while remaining competitive in visual understanding. Our work shows that unifying how AI sees and creates images under one framework is both feasible and beneficial.