BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Abstract
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual efficacy in real-world scenarios. To tackle this challenge, we present BizFinBench.v2, the first integrated offline and online benchmark built upon authentic user query-response data from both Chinese and U.S. equity markets. It comprises 28,860 questions across eight offline and two online tasks. Experimental results show that GPT-5 achieves a mere 61.5\% accuracy, still failing to meet the practical business requirement (84.8\%). Among the evaluated commercial models, DeepSeek-R1 exhibits superior investment efficacy. Error analysis grounded in real financial practice reveals persistent limitations in existing models. By overcoming the constraints of prior benchmarks, BizFinBench.v2 provides a substantiated foundation for advancing LLM deployment in the financial sector. Our data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.
Lay Summary
Large language models are widely used in finance, but most existing evaluation benchmarks rely on simulated or ordinary data. This causes model test results to differ greatly from their real performance in actual financial business. To fix this issue, we build BizFinBench.v2, a benchmark using real user data from Chinese and US stock markets. It includes 28,860 questions covering offline and online financial evaluation tasks. Our results find that even advanced large models cannot reach the accuracy standard required for real financial work, and mainstream models still have obvious practical flaws. BizFinBench.v2 offers a more realistic and reliable way to assess financial large language models. It also helps researchers and companies better understand model weaknesses and promotes the stable application of LLMs in real financial scenarios.