TRACER: Trust-Calibrated Offline-to-Online Reinforcement Learning
Yutong Zhang ⋅ Yaoran Yang
Abstract
Offline-to-online reinforcement learning (O2O RL) can reuse historical data, but stale or corrupted logs can also anchor online learning to the wrong model. We study a local alternative to fixed replay ratios: each time-state-action region receives a trust weight determined by offline coverage, online-offline reward/transition disagreement, and a confidence radius. The resulting algorithm, \tracer, uses this weight to mix empirical models, modulate optimism, and apply a local suspicion penalty. In finite-horizon tabular MDPs, we prove a high-probability interpolation bound with complete in-text proofs: clean covered regions contribute low-variance offline error, while identifiable shifted regions contribute an exponentially attenuated offline-bias term. Empirically, we evaluate 17 tabular algorithms and ablations on 18 controlled O2O tasks. These experiments should be read as a controlled mechanism study, not as a D4RL/MuJoCo or exact published-code comparison: ROAD-, ARB-, WSRL-, and RLPD-style baselines are tabular proxies. Within this scope, \tracer obtains the highest aggregate final return (82.86$\pm$3.19), best 10th-percentile return (65.1), and lowest failure rate (2.8\%). Regime-level results are mixed--\tracer is positive in 6 of 15 family/corruption cells and negative in all Layered cells against the strongest cell-wise proxy--which clarifies both the promise and current limits of local trust calibration.
Chat is not available.
Successful Page Load