Adaptive Estimation and Inference in Semi-parametric Heterogeneous Clustered Multitask Learning via Neyman Orthogonality
Abstract
We study clustered multitask learning in a semiparametric setting where tasks share a latent cluster structure in their target parameters but exhibit heterogeneous, potentially infinite-dimensional nuisance components. Such heterogeneity poses a major challenge for existing multitask learning methods, which typically rely on aligned feature spaces or homogeneous task structures. To address this challenge, we propose an adaptive fused orthogonal estimator that integrates Neyman-orthogonal losses with data-driven pairwise fusion penalties. Our framework leverages task-specific pilot estimates to calibrate the fusion penalties and combines adaptive aggregation with orthogonalization to mitigate the impact of nuisance-parameter estimation error. Theoretically, we show that the proposed estimator achieves exact recovery of the latent clustering with high probability and attains pooled parametric convergence rates proportional to cluster size. Moreover, we establish asymptotic normality and show that, asymptotically, our estimator matches the performance of an oracle procedure that knows the true clustering in advance. Empirically, we show that the proposed method consistently outperforms strong baselines in various simulation setups. A real-world application to U.S. residential energy consumption further demonstrates the effectiveness of our approach in uncovering meaningful regional clustering in electricity price elasticity, showcasing the efficacy of our method.
Lay Summary
Many scientific and policy problems require estimating related quantities across multiple datasets, such as treatment effects across hospitals or price sensitivities across states. An important challenge is deciding when these datasets should share information. Estimating each task separately can be unreliable/noisy when sample sizes are limited, whereas pooling all tasks together can be misleading when the parameter of interest differs across the tasks. This paper proposes a method that adaptively learns which tasks have similar target quantities and shares information only within those learned groups. At the same time, it allows each task to have completely heterogeneous nuisance parameters that are not particularly related to the parameter of interest. This helps prevent the sharing of invalid information while still improving statistical efficiency. We provide theoretical guarantees showing that our proposed method can recover the hidden groups of related tasks, estimate the target quantities at the same rate as if the groups were known in advance, and support valid statistical inference. In simulations, the method improves estimation accuracy and clustering performance compared with several competing methods. In an application to U.S. residential energy consumption, it identifies interpretable regional clusters in the elasticity of electricity prices.