MORE: A Multilingual Document Parsing Benchmark and Evaluation
Abstract
Lay Summary
(1) Problem: Documents are the ultimate vessels of human knowledge, and Artificial Intelligence is increasingly used to read them. However, most AI models are only tested on high-resource languages like English or Chinese. This creates a massive "blind spot": while modern AI claims to understand hundreds of languages, we actually have no reliable way to verify if it can accurately read a complex table in Welsh or a receipt in Amharic. (2) Solution: To fix this, we built MORE, the most diverse reading test for AI ever created. We collected real-world documents from the internet and carefully annotated them across 149 different languages. Unlike previous tests that only check simple text, MORE challenges AI to understand complex visual structures like tables, mathematical formulas, and computer code in these underrepresented languages. (3) Impact: By testing current state-of-the-art AI models on MORE, we discovered that while they are good at reading simple text, they still struggle heavily with complex layouts in less common scripts. Our benchmark provides a much-needed yardstick to help researchers build truly global, inclusive AI systems that leave no language behind.