Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer
Yannis Montreuil ⋅ Wei Tsang Ooi ⋅ Axel Carlier ⋅ Lai Xing Ng
Abstract
Existing multi-expert learning-to-defer surrogates are statistically consistent, yet they underfit under redundancy, suppress rare specialists, and degrade as the expert pool grows. We trace these failures to a shared architectural choice: casting classes and experts as actions inside one augmented prediction geometry. Consistency governs the population target; it says nothing about how the surrogate distributes gradient mass during training. We analyze five surrogates along both axes and show that each trades a fix on one for a failure on the other. We then introduce a decoupled surrogate that estimates the class posterior with a softmax and each expert utility with an independent sigmoid. It admits a consistency bound whose constant is $J$-independent for fixed per-expert weight $\lambda/J$, and its gradients are free of the amplification, starvation, and coupling pathologies of the augmented family. Experiments on synthetic benchmarks, CIFAR-10, CIFAR-10H, and Covertype confirm that the decoupled surrogate is the only method in our comparison that avoids amplification under redundancy and preserves rare specialists, while improving over a standalone classifier across the real-data benchmarks.
Chat is not available.
Successful Page Load