Mode-Aware Phenotype Profiling from Korean Clinical Reports: An LLM-Derived Two-Layer Fingerprint for Autism Characterization
Abstract
Autism spectrum disorder (ASD) is phenotypically heterogeneous. Clinical diagnostic reports provide a richer and more clinically proximal basis for heterogeneity analysis than questionnaire totals. We analyzed 309 Korean clinical reports with a sentence-attention KLUE-RoBERTa model and an LLM silver-labeling workflow that mapped high-attention clinical sentences into 18 DSM-5-grounded phenotype domains. Report-level domain frequency vectors agreed closely with expert-gold profiles. To test for discrete subtypes, we fit Gaussian mixture models on isometric log-ratio coordinates and benchmarked the resulting clusters against a permutation null with multi-seed stability checks. Four \emph{weak density modes} emerged consistently across random seeds but were not well separated enough for hard subtype assignment, comprising a dual prototype core, a broad diffuse mode, and a physiological-adaptive directional mode. Mahalanobis tail tests then located individual idiosyncrasy in the residual PC subspace rather than in the leading mode-carrying dimensions. We summarize the phenotype space with a two-layer fingerprint, mode membership plus within-mode tail deviation per domain, in place of hard subtype labels.