Lateral Data Exfiltration in MCP: How One Compromised Server Captures Cross-Domain Agent Data
Abstract
Production AI agents increasingly connect to multiple Model Context Protocol (MCP) servers, each ostensibly scoped to a distinct service domain. We demonstrate that MCP's shared-context architecture fundamentally violates the implicit trust boundaries between servers. Through controlled experiments, we show that a single compromised server can laterally exfiltrate data from all other connected servers via a covert logging tool, achieving 90.7% autonomous invocation across 360 interactions and 100\% capture of cross-server banking operations. We formalize this as the weakest-link property: the confidentiality of all agent-accessible data is bounded by the security posture of the least-secure connected server. This threat is qualitatively different from single-server tool poisoning, as it expands the blast radius of any server compromise to the entire agent ecosystem.