Adversarially Robust Approximate Furthest Neighbor
Kiarash Banihashem ⋅ Jeff Michael Giliberti ⋅ Prashant Gokhale ⋅ Samira Goudarzi ⋅ MohammadTaghi Hajiaghayi ⋅ Yuhao Liu ⋅ Morteza Monemizadeh ⋅ Sandeep Silwal
Abstract
We work in the adaptive query model, where one is given a point set $P \subset \mathbb{R}^d$ and seeks to construct a data structure that can answer correctly and efficiently a sequence of adaptive queries. In this model, an adversary observes the answers returned by the data structure to previous queries $q_1, \ldots, q_{i-1}$ and, based on this information, chooses the next query point $q_i$. This setting captures strong forms of adaptivity that naturally arise in modern machine learning pipelines, and rules out many classical randomized techniques that assume oblivious queries. Our focus is the problem of furthest neighbor search in this adaptive setting, a fundamental problem in several learning tasks, including diversity maximization, outlier and anomaly detection, adversarial example generation, and more. We present the first adversarially robust data structure for $c$-approximate furthest neighbor queries that achieves query time $\tilde{O}( \min( d n^{1/c^2}, n^{2/c^2} + d))$. This matches the $n$ dependency in the query time of the seminal result by Indyk [SODA'03] for $c$-approximate furthest neighbor in the oblivious setting, and improves upon the $\tilde{O}(n + d)$ query time achieved via the adaptive distance estimation framework of Cherapanamjeri and Nelson [NeurIPS'20] for a wide range of natural parameters. To complement this result, we present an adversarial attack against oblivious approximate furthest neighbor algorithms. Specifically, we show that the data structure from the algorithm by Indyk fails to maintain its guarantees against adaptive queries.
Lay Summary
Modern machine learning applications such as detecting data anomalies need to quickly find a data point that is "furthest" from a target. Traditional algorithms assume these search requests are independent and oblivious. In the real world, however, a clever adversary can study the system's past answers to strategically engineer new requests designed to trick the system into failing. We created an adversarially robust data structure for furthest-neighbor searches that remains both highly accurate and fast, even when an attacker is actively trying to break it. We also demonstrate how older, standard algorithms break under these adaptive attacks.
Successful Page Load