Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding
Abstract
Lay Summary
AI systems that answer questions about images can sometimes describe things that are not actually there, such as inventing objects or attributes. This makes them harder to trust in real-world uses where visual accuracy matters. Our work studies why these mistakes happen and finds that the model is not simply “looking too little” at the image. Instead, some parts of the model can rely too strongly on learned language habits, especially when deciding what to say next. We propose Fox, a method that detects these risky parts during generation and reduces their influence, without retraining the model. The method also keeps the model’s ability to produce detailed and natural answers, so it does not become overly cautious. In tests across several image-language AI systems, Fox reduces false visual claims while keeping answers useful and fluent, with little extra cost. This work offers a practical step toward more reliable AI systems that can describe and reason about visual content more faithfully.