Keynote #3: From Scores to Process Evidence: Automated Information Triage for Long-Horizon Agent Evaluations
Abstract
Long-horizon agent evaluations are increasingly important for safety and capability assessment, but they strain current benchmarking practice. Runs are costly, sample sizes are small, and transcripts are too large for consistent expert review. Final outcomes may also miss key evidence about progress, failure modes, shortcut use, and harness effects.
I present Automated Information Triage (AIT), a method and review interface for turning long agent trajectories into expert reviewer-facing process evidence. AIT starts with expert calibration of the evaluation claim, success criteria, milestones, artefacts, and review rubrics. It then uses transcript scanners, artefact analysis, and LLM-judge enrichment to produce key decision points, run narratives, process signals, and source-linked evidence.
Using insights from an AI R&D case study, I discuss open questions around signal validity, transfer across tasks, and the boundary between reviewer assistance and automated judgement.