Tool-Existence Hallucination Is Tier-Dependent and Intra-Family Non-Monotonic on BFCL Multi-Turn
S Akash ⋅ Shine Gupta
Abstract
We define per-call Tool-Existence Hallucination Rate (TEHR) and run it over BFCL multi-turn against two populations: five Anthropic 4.x versions spanning an eleven-month release window, and the Qwen3-Instruct family at six sizes from 0.6B to 32B (4-bit MLX). The Anthropic side is quiet. Across 2,599 calls under an opaque baseline we log zero TEHR events, with a Clopper--Pearson 95\% upper bound of $\leq 0.115\%$. Qwen3 is not. Its rate is non-monotone in scale, rising from $0\%$ at 0.6B to $1.87\%$ at 14B and falling back to $0\%$ at 32B, while the dominant distractor type slides from near-name (1.7B) to matched-random (4B/8B) to synonym (14B). We propose Registry-Visible Reprompting (RVR), a training-free middleware that checks each proposed call against the runtime registry and, on a miss, returns a single re-prompt wrapping a structured \texttt{tool\_not\_found} envelope. RVR removes every one of the fourteen pooled Qwen3 fabrications (Fisher's exact one-sided, $p = 7.1 \times 10^{-5}$ on $14/973$ vs.\ $0/945$). A content-matched ablation that drops the registry list (C0.7; $0/253$ on Qwen3-8B, matching C1) localizes the effect to the envelope shape rather than to listing the actual tool names. Deployers can therefore ship the cheaper signal without echoing internal registry contents. On the Anthropic zero-event regime RVR also logs zero events, but the strict-pass subset slips at $N = 60$, so we recommend tier-conditional deployment. We argue tool-name hallucination here is a mid-scale Qwen3 phenomenon. A one-line structured reply is enough to remove it on the models where it appears.
Chat is not available.
Successful Page Load