Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory
Abstract
Shadow trading—trading in a peer firm’s securities on the basis of material nonpublic information (MNPI) about an “economically linked” company—is a novel and contested theory of insider trading liability, first prosecuted in SEC v. Panuwat (2023). Enforcing it requires identifying economically linked firms ex ante, a determination the SEC makes only after the fact using mass market surveillance infrastructure. We ask whether NLP can do what the SEC’s theory presumes insiders already know: identify peer firms ex ante from publicly mandated disclosures. Using a two-stage LLM pipeline applied to Item 7 (Management’s Discussion and Analysis) sections of SEC 10-K filings, we score semantic similarity across 30 M&A events spanning five industries and correlate similarity ranks with announcement-day abnormal stock returns. Our pipeline replicates the Panuwat precedent on its most analogous case, validating the methodology. Across the full dataset, however, the relationship is indistinguishable from a coin flip: 14 of 30 events support the shadow trading hypothesis, 12 contradict it, and 4 are ambiguous (binomial p ≈ 0.86 against H_0 = 0.5). These results challenge the empirical premise of shadow trading enforcement and bear directly on constitutional questions surrounding the SEC’s mass financial surveillance infrastructure.