What Can Passive API Access Verify About Model Provenance?
Jeewoo Kim ⋅ Su Hyeong Lee
Abstract
Can an external auditor determine whether a deployed API is serving a derivative of a known model, using only standard API access, with no prompt control and no injected fingerprints? We show that passive API evidence can support provenance screening but not definitive attribution, and that a single returned log-probability per token is the lowest-burden transparency threshold we tested that enables such screening. Across 182 evaluated model variants in six families, a training-free sequence-likelihood gap achieves AUROC 0.85-1.00 for descendant detection against same-family hard negatives. However, the same signal cannot identify which specific derivative is deployed, recover lineage trees, or survive adversarial system prompts. Without logprob access, this screening signal collapses entirely, though text-similarity methods remain viable when prompts are known. These results map the auditability frontier of passive provenance: $k=1$ logprob return is a concrete, low-burden transparency standard that enables this class of third-party screening, but passive evidence should be treated as triage, not as enforcement-grade proof.
Chat is not available.
Successful Page Load