SEA-MU: Cultural Meme Understanding Benchmark for Southeast Asia
Abstract
Open-ended Visual Question Answering (VQA) is a multimodal task that requires both understanding and reasoning from visual and language information. Since existing Vision Language Models (VLMs) are pretrained dominantly on Western languages, the cultural knowledge of VLMs often skews toward Western cultures. Furthermore, existing cultural VQA benchmarks focus on simply recalling facts. These benchmarks fail to involve reasoning in cultural common sense, especially in the context of Southeast Asia (SEA). Meme, being a common form of Internet communication requiring the ability to identify not only the textual or visual cues but also the background knowledge, represents an interesting medium to evaluate VLM reasoning ability. In this ongoing work, we propose SEA-MU, a cultural meme understanding benchmark to evaluate the models' cultural reasoning ability in Thai, Vietnamese, and Indonesian contexts. Our preliminary experiments with annotated Thai memes suggest that current open-source VLMs struggle to reason effectively in SEA context.