Posterior-Driven Actor-Critic Framework for Active Hypothesis Testing
Abstract
We propose a readily adaptable, general purpose framework for learning active hypothesis testing policies across a wide variety of problem settings. Our PostAC framework trains a sequential, adaptive policy which tracks a Bayesian posterior over hypotheses and, at each time-step, uses a deep neural network to calculate the policy’s actions as a function of this posterior. We first lay out an actor-critic algorithm that efficiently trains these DNN policies and then apply our algorithm to two problems: channel coding with feedback, where we match the performance of an analytically derived stateof- the-art coding scheme, and the problem of coherent state discrimination for optical communication, where we show stateof- the-art performance in the low power regime.