MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
Abstract
Recently, vision-language models have demonstrated increasing influence in morally sensitive domains such as autonomous driving and medical analysis, owing to their powerful multimodal reasoning capabilities. As these models are deployed in high-stakes real-world applications, it is of paramount importance to ensure that their outputs align with human moral values and remain within moral boundaries. However, existing work on moral alignment either focuses solely on textual modalities or relies heavily on AI-generated images, leading to distributional biases and reduced realism. To overcome these limitations, we introduce MORALISE, a comprehensive benchmark for evaluating the \underline{mor}al \underline{al}ignment of v\underline{is}ion-languag\underline{e} models (VLMs) using diverse, expert-verified real-world data. We begin by proposing a comprehensive taxonomy of 13 moral topics grounded in Turiel's Domain Theory, spanning the personal, interpersonal, and societal moral domains encountered in everyday life. Built on this framework, we manually curate 2,481 high-quality image-text pairs, each annotated with two fine-grained labels: (1) \textit{topic annotation}, identifying the violated moral topic(s), and (2) \textit{modality annotation}, indicating whether the violation arises from the image or the text. For evaluation, we encompass two tasks, \textit{moral judgment} and \textit{moral norm attribution}, to assess models' awareness of moral violations and their reasoning ability on morally salient content. Extensive experiments on 19 popular open- and closed-source VLMs show that MORALISE poses a significant challenge, revealing persistent moral limitations in current state-of-the-art models.
Lay Summary
AI systems that look at images and read text are increasingly used in areas such as driving, healthcare, and education, where a wrong judgment can affect people. But it is still hard to know whether these systems understand everyday moral problems, especially when the key clue appears in a photo rather than in the words. We built MORALISE, a test collection designed to measure how well vision-language AI systems recognize and explain moral violations. It contains 2,481 real-world image–text examples, checked by human annotators, and covers 13 types of moral concerns such as harm, fairness, responsibility, respect, and discrimination. Unlike many earlier tests, MORALISE uses real images and marks whether the moral issue comes mainly from the image or the text. We tested 19 AI models on two questions: whether something is morally wrong, and which moral concern is being violated. The models often did reasonably well at the first question, but struggled much more with the second, especially when the relevant evidence was visual. These findings show that multimodal AI systems still need better ways to understand moral situations before they are used in sensitive real-world settings.