Reinforcement Learning for Adaptive Tacrolimus Dosing with Multi-Drug Interaction Management
Abstract
Tacrolimus is a critical immunosuppressant following solid organ transplantation, but its narrow therapeutic window (4–15 ng/mL) and high interpatient pharmacokinetic variability make dosing difficult, particularly when CYP3A4-inhibiting co-medications cause trough concentrations to rise. We formulate tacrolimus dosing with multidrug drug-drug interaction (DDI) management as a Partially Observable Markov Decision Process (POMDP) and train an LSTM-augmented Proximal Policy Optimization (LSTM-PPO) model in a two-compartment physiologically-based pharmacokinetic (PBPK) simulator spanning four clinically relevant CYP3A4 inhibitors. The agent achieves 93.1% time in therapeutic window (TITW), a 13.9 percentage-point improvement over a proportional TDM controller. Against a memoryless ablation, the agent reduces toxic concentration events 8-fold and rejection risk sixfold, demonstrating that memory specifically prevents tail failures in DDI-exposed patients. Retrospective EHR-based validation on 6,394 real patients from the Cedars-Sinai OMOP cohort demonstrates a 5.8 percentage-point higher simulated TITW than observed EHR trough trajectories, consistent across all demographic and clinical subgroups. These results support memoryaugmented RL as a promising framework for precision dosing from longitudinal, partially observed structured EHR data.