Beyond Autonomy: Human-AI Collaboration in Science
Abstract
Recent advances in LLMs have sparked excitement about AI for science, with growing evidence that these systems can generate ideas, run experiments, and accelerate discovery. Yet this framing centers the AI half of the loop, leaving the human half—who collaborates, who benefits, and how—underexamined. In this talk, we argue that realizing AI for science requires studying not just what models can do, but how people work with them. We first examine the ideation-execution gap: although LLM-generated ideas appear more novel than expert ideas at the proposal stage, our large execution study shows this advantage collapses once ideas are implemented, with AI scores dropping significantly. We then ask whether grounding research in real execution can close this gap, presenting a system that builds automated executors and learns from large-scale experimental feedback via evolutionary search and reinforcement learning. Finally, we ask how to measure collaboration itself, presenting CollabSkill, a framework that uses Bayesian skill rating to disentangle human and agent contributions and reveals that collaboration rankings diverge from autonomous benchmarks. We conclude by discussing how realizing the promise of AI for science depends on not only model capability, but also on designing the collaboration between them.