Towards Automated Evaluation of Socio-Technical Harms in LLMs: A Normative Taxonomy and Multi-Turn Red-Teaming Framework
Abstract
Evaluating the safety of large language models (LLMs) with respect to socio-technical harms requires more principled criteria than surface-level toxicity or task performance metrics without clear normative grounding. This paper presents an interdisciplinary framework, developed with philosophers, legal scholars, and social scientists, for the automated evaluation of socio-technical harms in LLMs. We construct a normative taxonomy of harmful LLM behaviors across four categories, grounded in international AI governance instruments and moral-philosophical theory, and empirically calibrated through a lay-user survey with 250 participants. We then instantiate this taxonomy in an automated evaluation pipeline in which a Strategic Dialogue Actor (attacker LLM) generates multi-turn adversarial conversations using a phase-adaptive Observation–Thought–Strategy–Reply (OTSR) architecture, and an LLM-as-a-judge (judge LLM) scores each turn using category-specific rubrics derived directly from the taxonomy. This framework enables fine-grained, scalable assessment of whether LLMs maintain normative compliance across harm categories, adversarial techniques, and realistic conversational pressures.