RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
Abstract
Large Language Models (LLMs) are increasingly vulnerable to Prompt Injection (PI) attacks, where adversarial instructions hidden within retrieved contexts hijack the model's execution flow. Current defenses typically face a critical trade-off: prevention-based fine-tuning often degrades general utility via the "alignment tax", while detection-based filtering incurs prohibitive latency and memory costs. To bridge this gap, we propose RedVisor, a unified framework that synthesizes the explainability of detection systems with the seamless integration of prevention strategies. To the best of our knowledge, RedVisor is the first approach to leverage fine-grained reasoning paths to simultaneously detect attacks and guide the model's safe response. We implement this via a lightweight, removable adapter positioned atop the frozen backbone. This adapter serves a dual function: it first generates an explainable analysis that precisely localizes the injection and articulates the threat, which then explicitly conditions the model to reject the malicious command. Uniquely, the adapter is active only during this reasoning phase and is effectively muted during the subsequent response generation. This architecture yields two distinct advantages: (1) it mathematically preserves the backbone's original utility on benign inputs; and (2) it enables a novel KV Cache Reuse strategy, eliminating the redundant prefill computation inherent to decoupled pipelines. We further pioneer the integration of this defense into the vLLM serving engine with custom kernels. Experiments demonstrate that RedVisor outperforms state-of-the-art defenses in detection accuracy and throughput while incurring negligible utility loss.
Lay Summary
Artificial intelligence assistants are often asked to read large amounts of information, making them vulnerable to hidden "prompt injection" attacks. These malicious instructions can trick the AI into breaking its safety rules. Current defenses force a difficult choice: they either permanently alter the AI, making it worse at answering normal questions, or they use slow, separate safety filters. To solve this, we developed RedVisor, a lightweight safety shield that attaches directly to the AI. Before answering a user, RedVisor inspects the text and explicitly points out any hidden attacks. After this security check, the shield temporarily mutes itself, allowing the original, unchanged AI to safely respond. Because the AI and the shield share the same memory, the system skips redundant reading steps, making it highly efficient compared to older multi-step methods. Our findings show that RedVisor effectively blocks sophisticated attacks while maintaining the AI's original utility and keeping response times low. This approach provides a practical way for developers to build safer, more reliable AI tools for real-world applications without severe performance trade-offs.