Abstract:Weakly supervised whole slide image classification is a key task in computational pathology, which involves predicting a slide-level label from a set of image patches constituting the slide. Constructing models to solve this task involves multiple design choices, often made without robust empirical or conclusive theoretical justification. To address this, we conduct a comprehensive benchmarking of feature extractors to answer three critical questions: 1) Is stain normalisation still a necessary preprocessing step? 2) Which feature extractors are best for downstream slide-level classification? 3) How does magnification affect downstream performance? Our study constitutes the most comprehensive evaluation of publicly available pathology feature extractors to date, involving more than 10,000 training runs across 14 feature extractors, 9 tasks, 5 datasets, 3 downstream architectures, 2 levels of magnification, and various preprocessing setups. Our findings challenge existing assumptions: 1) We observe empirically, and by analysing the latent space, that skipping stain normalisation and image augmentations does not degrade performance, while significantly reducing memory and computational demands. 2) We develop a novel evaluation metric to compare relative downstream performance, and show that the choice of feature extractor is the most consequential factor for downstream performance. 3) We find that lower-magnification slides are sufficient for accurate slide-level classification. Contrary to previous patch-level benchmarking studies, our approach emphasises clinical relevance by focusing on slide-level biomarker prediction tasks in a weakly supervised setting with external validation cohorts. Our findings stand to streamline digital pathology workflows by minimising preprocessing needs and informing the selection of feature extractors.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the impact of the choice of feature extractors and their pre - processing steps (such as stain normalization and image augmentation) on the model performance in the weakly - supervised whole - slide - image (WSI) classification task. Specifically, the paper answers the following three key questions through comprehensive benchmark tests: 1. **Is stain normalization still a necessary pre - processing step?** - The paper explores whether stain normalization is still necessary when using self - supervised - learning (SSL) feature extractors. In the traditional digital pathology workflow, stain normalization is a standard pre - processing step used to correct image differences caused by different scanners and stains. However, with the development of SSL feature extractors, the necessity of this step is worth re - evaluating. 2. **Which feature extractors are most suitable for downstream slide - level classification tasks?** - The paper compares a variety of publicly available pathology feature extractors, aiming to determine which feature extractors perform best in downstream tasks. These feature extractors include models pre - trained on ImageNet and models self - supervised - trained on large - scale pathology datasets. 3. **How does magnification affect downstream performance?** - The paper studies the model performance at different magnifications (low magnification and high magnification) to determine which magnification is more suitable for accurate slide - level classification. To answer these questions, the paper conducted more than 10,000 training runs, covering 14 feature extractors, 9 tasks, 5 datasets, 3 downstream architectures, 2 magnifications, and different pre - processing settings. The research results challenge existing assumptions, and the main findings are as follows: 1. **Skipping stain normalization and image augmentation does not reduce performance** while significantly reducing memory and computational requirements. 2. **The choice of feature extractor is the most critical factor affecting downstream performance**. The paper developed a new evaluation metric to compare relative downstream performance. 3. **Low - magnification slides are sufficient for accurate slide - level classification**. These findings are expected to simplify the digital pathology workflow, reduce pre - processing requirements, and guide the choice of feature extractors.

Benchmarking Pathology Feature Extractors for Whole Slide Image Classification

Benchmarking Embedding Aggregation Methods in Computational Pathology: A Clinical Data Perspective

The Importance of Downstream Networks in Digital Pathology Foundation Models

Benchmarking foundation models as feature extractors for weakly-supervised computational pathology

Overcoming the limitations of patch-based learning to detect cancer in whole slide images

Exploring Pathologist Knowledge for Automatic Assessment of Breast Cancer Metastases in Whole-slide Image.

Clinical-grade computational pathology using weakly supervised deep learning on whole slide images

Augmenting the Pathology Lab: An Intelligent Whole Slide Image Classification System for the Real World

Developing image analysis pipelines of whole-slide images: Pre- and post-processing

The Whole Pathological Slide Classification via Weakly Supervised Learning

PATHS: A Hierarchical Transformer for Efficient Whole Slide Image Analysis

Optimization of whole slide imaging scan settings for computer vision using human lung cancer tissue

Low-resource finetuning of foundation models beats state-of-the-art in histopathology

Classification in Histopathology: A unique deep embeddings extractor for multiple classification tasks

Examining Batch Effect in Histopathology as a Distributionally Robust Optimization Problem

An attention-based multi-resolution model for prostate whole slide imageclassification and localization

Automatic Whole Slide Pathology Image Diagnosis Framework Via Unit Stochastic Selection and Attention Fusion

Scalable deep learning artificial intelligence histopathology slide analysis and validation

A whole-slide foundation model for digital pathology from real-world data

Contrastive learning-based histopathological features infer molecular subtypes and clinical outcomes of breast cancer from unannotated whole slide images

Efficient Classification of Histopathology Images