Self-correcting for Debiasing Large Language Models
Abstract
Although Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, inherent social biases often cascade throughout the Chain-of-Thought (CoT) process, leading to continuous "Bias Propagation". Existing debiasing methods primarily focus on static constraints or external interventions, failing to identify and interrupt this propagation once triggered. To address this limitation, we introduce Self-Debias, a progressive framework designed to instill intrinsic self-correction capabilities. Specifically, we reformulate the debiasing process as a strategic resource redistribution problem, treating the model's output probability mass as a limited resource to be reallocated from biased heuristics to unbiased reasoning paths. Unlike standard preference optimization which applies broad penalties, Self-Debias employs a fine-grained trajectory-level objective subject to dynamic debiasing constraints. This enables the model to selectively revise biased reasoning suffixes while preserving valid contextual prefixes. Furthermore, we integrate an online self-improvement mechanism utilizing consistency filtering to autonomously synthesize supervision signals. With merely 20k annotated samples, Self-Debias activates efficient self-correction, achieving superior debiasing performance while preserving general reasoning capabilities without continuous external oversight.
Lay Summary
Large Language Models often solve problems by generating step-by-step reasoning. However, if social biases appear in the early reasoning steps, they may spread through the following steps and lead to unfair or stereotyped answers. This paper aims to help models detect and correct such biased reasoning by themselves. We propose Self-Debias, a training framework that teaches models to revise only the biased parts of their reasoning while keeping the useful and correct parts unchanged. Instead of simply penalizing an entire response, Self-Debias focuses on where biased reasoning begins and encourages the model to shift toward fairer and more evidence-based reasoning paths. The framework also allows the model to improve itself by generating and selecting additional useful training examples. With only 20,000 annotated samples, Self-Debias reduces biased outputs while preserving general reasoning ability. This work provides a step toward language models that can better recognize, interrupt, and correct bias during their own reasoning process.