BFCL Audio: An Audio Function Calling Evaluation for Large Language Models
Abstract
Lay Summary
When you ask a voice assistant to do something ("book a flight to Austin"), it has to turn your spoken words into an exact command. But speech is messy — accents, stumbles, background noise — so the assistant can mishear a detail and confidently do the wrong thing (book "Boston" instead). Until now there was no good way to test how often this happens or why. The authors built BFCL Audio: a kit of ~6,200 tasks that deliberately adds realistic noise, accents, and stumbles to measure how reliably voice assistants carry out spoken requests, and grades them automatically. Key finding: noise makes every system noticeably worse (even the best dropped ~10%), and the worst mistakes are quiet ones — mishearing names or numbers, or babbling instead of acting. The test kit is being released publicly so others can build more trustworthy voice assistants.