ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
Abstract
Lay Summary
This paper introduces ToolOrchestra, a method for training a small "manager" AI that, instead of solving hard problems on its own, decides which other AI models and tools to call on and coordinates them to reach an answer. Trained with reinforcement learning that rewards correct answers, low cost, and respect for user preferences about which tools to use, the resulting 8-billion-parameter "Orchestrator" punches well above its weight: on the notoriously difficult Humanity's Last Exam benchmark it scored 37.1%, beating the far larger GPT-5 (35.1%) while being about 2.5 times more cost-efficient, and on other benchmarks it outperformed GPT-5 by a wide margin at roughly a third of the cost while also adapting well to tools it had never seen—suggesting that a lightweight coordinator paired with diverse specialized tools can be both cheaper and more capable than relying on a single enormous model.