Are We Overconfident in Models and Results for Semi-Supervised 3D Medical Image Segmentation?
Abstract
Semi-supervised learning has become a dominant paradigm for reducing annotation costs. However, we argue that the current progress is clouded by a twofold overconfidence problem. Algorithmically, mainstream pseudo-labeling frameworks often conflate prediction confidence with uncertainty, leading to severe confirmation bias. Strategically, since multiple benchmark datasets lack dedicated validation sets, some studies use the test set for validation as well, leading to inflated performance estimates. Subsequent methods, compelled to employ the same strategy to surpass reported SOTA, trigger an arms race of overfitting. This raises concerns that the impressive numerical gains in the community may reflect overfitting rather than genuine progress. Thus, we propose a tri-space calibrated segmentation framework founded on a principled dual-axis reliability assessment engine. It explicitly decouples confidence from uncertainty and uses this signal to detect and correct confirmation bias across feature, probability, and image spaces in a collaborative manner. Across three benchmark datasets, TCSeg consistently delivers strong performance under existing evaluation protocols. More importantly, we advocate that the community report final-checkpoint results under multiple-run protocols, thereby establishing more rigorous benchmarks with a more realistic perspective. Code will be available: github.com/DirkLiii/TCSeg.
Lay Summary
Semi-supervised learning could make medical AI cheaper and more practical by reducing the need for experts to label every medical scan. But current progress may be overly optimistic. Many methods trust their own automatically generated labels too much, so early mistakes can be reinforced during training. In addition, some common evaluation practices can make performance look better than it really is. We propose a new training method that is more careful about which machine-generated labels to trust. It checks predictions from multiple perspectives and uses this information to improve learning in several complementary ways. Tests on three widely used medical imaging benchmarks show that our method is both accurate and stable. Just as importantly, we argue that the field should use stricter evaluation practices, such as repeated experiments and reporting the final trained model, so that progress is measured more realistically. This could help make future medical AI systems more dependable for real-world use.