Deep Learning for BioImaging: What Are We Really Learning?
Abstract
Representation learning has driven major advances in natural image analysis by enabling models to acquire high-level semantic features. In microscopy imaging, however, it remains unclear what current representation learning methods really learn. In this work, we conduct a systematic study of representation learning for the two most widely used and broadly available microscopy data types, representing critical scales in biology: cell culture and tissue imaging. We investigate whether, in contrast to natural images, existing models fail to consistently acquire high-level, biologically meaningful features. To this end, we introduce a set of simple yet revealing baselines on curated benchmarks, including untrained models and structural representations of cellular tissue. Our results show that, surprisingly, for a considerable subset of evaluation settings, the baselines are comparable to state-of-the-art methods, demonstrating that many commonly used benchmark metrics are insufficient to assess representation quality and often mask a lack of relevant high-level abstractions. In addition, we investigate how detailed comparisons with these baselines provide ways to interpret the strengths and weaknesses of models for further improvements. Together, our results suggest that progress in representation learning for microscopy requires not only stronger models, but also benchmarks that are more indicative of what is actually learned.
Lay Summary
Benchmarks are model problems that help researchers evaluate and compare the performance of AI models. Ideally, we want models that learn something practically useful to perform better than the rest. However, this is often not the case. In our work, we focus on highlighting this issue for benchmarks in microscopy imaging of human cells. We achieve this by designing naive approaches that have no relevant biological knowledge. We then show how benchmarks that were designed to rank the most powerful domain-specific AI models often fail to distinguish them from our naive tools. Addressing this limitation is critical for biomedical research. Better evaluation leads to better understanding, better training, and, eventually, a safer integration of AI into medical diagnostics and scientific discovery.