Empirical-Distribution Matching for Synthetic ECG Classification
Benedikt Kolbeinsson ⋅ Arinbjörn Kolbeinsson
Abstract
Conditional generative models are increasingly proposed as drop-in substitutes for restricted clinical datasets: train a generator once, release a synthetic cohort, and let practitioners reuse it without ever touching the source records. The practitioner's only design surface is then the sampling policy used to draw the synthetic training set. The main idea of this paper is that the standard rebalancing recipes from the CV and NLP imbalanced-learning literature do not transfer to this regime; the practitioner is better off matching the empirical training distribution as faithfully as possible at sample time. We test the idea by surveying thirteen sampling policies on a single $12$-lead ECG latent diffusion model over PTB-XL using train-on-synthetic-test-on-real (TSTR) macro AUROC. No policy beats the naive bootstrap baseline at matched budget; qstrat_matched ties at one seed. The failures collapse onto three mechanisms: class-distribution distortion, within-class diversity collapse, and label-by-demographic joint decoupling.
Chat is not available.
Successful Page Load