Flow-Based Offline Reinforcement Learning for Voltage Regulation in Distribution Networks
Abstract
This paper investigates pure data-driven active voltage control in distribution networks via offline reinforcement learning (RL) to minimize the risks of online interactions. To overcome the limited policy expressivity of existing algorithms on low-quality, randomly collected datasets, we propose an improved Flow Q-Learning (improved FQL) approach featuring a Flow-Guided Base Action and Fine-Tuning Perturbation framework. This architecture utilizes a flow-matching model to accurately capture the distribution of safe behaviors, while a residual actor applies bounded perturbations to preserve physical safety while maximizing Q-values. Furthermore, Boltzmann annealing and Zero-Noise Initialization mechanisms are introduced to enhance convergence and execution stability. Simulations on IEEE 33-bus and 123-bus systems demonstrate that our approach outperforms baseline offline RL algorithms in mitigating voltage deviations and violations, despite relying entirely on randomly collected offline experience.