LCIA Self-Play: The London Court of International Arbitration as a Benchmark for Verifiable Self-Evolving Scientific Agents
Abstract
Self-evolving scientific agents are often evaluated where the training record is observable and correctness can be checked directly, but many scientific and institutional settings expose only selected public traces, contain strategic actors, and impose formal admissibility constraints. LCIA Self-Play is a synthetic benchmark inspired by the London Court of International Arbitration that models confidential arbitration as a partially observable hidden-type game with a deterministic legality gate over procedural actions. Agents improve through self-play, belief updating, episodic memory, and capability-preserving rehearsal while being evaluated on calibration, public-slice overfitting, hidden-type robustness, exploitability, violation probability, and retention. Across twelve pre-specified seeds, public-slice supervision preserves ranking signal but is badly miscalibrated on the hidden population (Brier 0.238 +/- 0.001 to 0.254 +/- 0.002; ECE 0.017 +/- 0.002 to 0.120 +/- 0.006), while verifier-gated self-play raises worst-hidden-type value over public-greedy play (0.626 +/- 0.007 versus 0.510 +/- 0.017) with zero invalid-action mass. Memory replay improves retained old-family value over recent-only adaptation (0.685 +/- 0.008 versus 0.656 +/- 0.009), making the contribution a reproducible stress test for constrained self-evolution rather than a legal prediction system.