Code-Switching Reveals Anchor Bias in Multilingual Large Language Models
Abstract
Multilingual Large Language Models (MLLMs) are increasingly expected to handle inputs with multiple languages across diverse contexts. However, it remains unclear how MLLMs internally organize code-switching (CS), a multilingual practice where speakers alternate between languages within a single interaction. In this paper, we use CS Question Answering (QA) as a controlled setting to probe language anchoring in MLLMs. Following common English--target-language evaluation settings, we treat English as the source and the paired non-English language as the target. We introduce anchor bias, a representation-level measure that assesses whether a CS hidden state is closer to its source or target-language counterpart. Across various MLLMs, we observe a consistent pattern: source-framed CS is more strongly anchored toward the source side, whereas target-framed CS is anchored toward the target side. This representational preference is reflected in QA behavior, where source-anchored inputs suffer smaller performance drops. We also verify that in MoE models, this anchoring appears not only in hidden-state geometry but also in sparse computational pathways. By quantifying how grammatical frames dominate internal organization, our work provides an interpretive framework for evaluating the linguistic sensitivity of MLLMs in mixed-language settings.