Unlocking Cross-Modal Biosignal Synthesis: A Temporally-Aware VAE-Diffusion Model
Abstract
Lay Summary
Heart monitoring often relies on two types of signals: electrical signals from an electrocardiogram (ECG) and mechanical heart sounds from a phonocardiogram (PCG). While ECGs are easy to capture with common devices, PCGs usually require specialized sensors, limiting access for many people. We asked: can we generate realistic heart sounds directly from widely available ECG recordings? To solve this, we designed a hybrid AI model that combines two powerful techniques. First, it learns a simplified representation of real heart sounds, capturing the essential patterns. Then, it uses this representation to generate high-quality heart sounds that match the timing and rhythm of the ECG. Our model explicitly learns how electrical signals drive mechanical heart activity, producing accurate waveforms even on previously unseen datasets. This approach could make dual-modality heart monitoring more accessible, supporting research, algorithm development, and data augmentation without needing extra hardware. While the generated sounds are not yet approved for medical diagnosis, they reliably mimic real heart signals and could help advance low-cost cardiac assessment and AI-assisted heart studies.