Robust ECG Classification under Patient-Wise Evaluation: A Study of Dynamical and Deep Representations
Abstract
Our work highlights a fundamental tension in ECG representation learning between predictive performance and cross-patient generalization. While deep architectures achieve strong results under patient-wise evaluation, their sensitivity to inter-subject variability raises concerns about robustness in real-world clinical deployment. Dynamical representations based on Koopman theory provide a complementary perspective by capturing global temporal structure in a linearized latent space, yielding more stable behavior across patients despite lower standalone performance. Their integration with deep models consistently improves robustness, indicating that structured dynamical features capture information not fully exploited by purely data-driven approaches. These findings suggest a broader design principle for physiological time-series modeling: combining data-driven learning with structured inductive biases can improve both performance and generalization. Hybrid approaches that integrate dynamical constraints with deep architectures therefore represent a promising direction for robust clinical AI. More broadly, our results emphasize the importance of evaluation protocols. Patient-wise splits are essential for realistic assessment, as standard splits can lead to overly optimistic estimates; robustness under distribution shift should thus be treated as a primary evaluation criterion. For controlled and reproducible comparison, we adopt a simplified single-label formulation based on keyword mapping, which abstracts away the inherently multi-label nature of ECG diagnoses (see Appendix~F). Future work will explore multi-label formulations, learned Koopman embeddings, subject-invariant representations, and validation on larger, more diverse clinical datasets.