Skip to yearly menu bar Skip to main content


Static Benchmarks Are Broken: The Case for Dynamic Evaluation of LLMs

Farhan Ahmed ⋅ Chad DeLuca

Abstract

Chat is not available.