Publications

Persistent Homology and Gabor Features Reveal Inconsistencies Between Widely Used Colorectal Cancer Training and Testing Datasets

D. Brito-Pacheco, Riad Ibadulla, X. Fernández, P. Giannopoulos, Constantino Carlos Reyes-Aldasoro

Medical Image Understanding and Analysis (MIUA 2025), Springer, 2025

In brief

Two colorectal cancer histology datasets, NCT-CRC-HE-100K and CRC-VAL-HE-7K, are widely used together as training and testing sets. Features extracted with persistent homology and Gabor filters show that the two sets have noticeably different distributions, which matters for anyone benchmarking models on them.

Overview

Research in computer vision and image processing relies heavily on open datasets, which allow techniques to be compared objectively. This paper examines two datasets that are commonly used together in colorectal cancer research: NCT-CRC-HE-100K, with 100,000 patches of haematoxylin and eosin stained tissue used for training, and CRC-VAL-HE-7K, with 7,180 patches used for testing.

Features extracted from both sets, first with persistent homology and then with Gabor filters, show that the training set has a noticeably different distribution from the testing set.

The paper was presented at MIUA 2025, the 29th Annual Conference on Medical Image Understanding and Analysis, in Leeds.

Citation

D. Brito-Pacheco, Riad Ibadulla, X. Fernández, P. Giannopoulos, Constantino Carlos Reyes-Aldasoro. “Persistent Homology and Gabor Features Reveal Inconsistencies Between Widely Used Colorectal Cancer Training and Testing Datasets.” Medical Image Understanding and Analysis (MIUA 2025), Springer, 2025. doi:10.1007/978-3-031-98688-8_7