RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
Yingsheng Geng ⋅ Yuchong Gao ⋅ Weihong Wu ⋅ Guyue Liu ⋅ Jiang liu
Abstract
The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, which significantly increases KV cache memory usage and time-to-first-token (TTFT). While various KV cache methods have been proposed to mitigate prefill redundancy, they either fail to maintain accuracy on agent-generated outputs or exhibit low reuse rates due to rigid constraints. We present RelayCaching, a training-free inference method that directly reuses decoding phase KV caches from previous agents in subsequent prefill phases. Our key insight is that KV caches for identical content are highly consistent across phases, while prefix-induced deviations are sparse and localized within a limited range of layers and token positions. By selectively recomputing KV caches at these positions, RelayCaching preserves model accuracy with minimal overhead, yielding a superior accuracy–efficiency trade-off over existing methods. Experiments on diverse collaborative LLM tasks spanning mathematical reasoning, general knowledge, and code generation demonstrate that RelayCaching achieves over $80$\% KV cache reuse, reduces TTFT by up to $4.7\times$ compared to the standard pipeline, all with negligible accuracy degradation.
Lay Summary
When multiple AI agents collaborate on complex tasks like writing research papers or solving scientific problems, each one must reprocess all the text that previous agents have already produced, creating redundant computation that grows with every additional participant. We introduce RelayCaching, a method that lets each agent directly reuse the internal computations of its predecessor. Like passing a relay baton, this avoids re-running the entire race. Since the reused results are not perfectly aligned with what the new agent expects, our method identifies and corrects only the small fraction of mismatched positions. This makes multi-agent systems much faster and more resource-efficient while preserving answer quality.
Successful Page Load