Health AIDeep learning, interpretability

Breast Cancer Histopathology Classification

A multi-magnification deep learning pipeline that classifies breast tissue images as benign or malignant, with Grad-CAM showing the evidence behind every prediction.

Test accuracy for 40X, 200X, and fusion
98.94%malignant recall
96.38%accuracy
97.40%F1 score
9,109images, 82 patients

The problem

Most histopathology classifiers look at one magnification at a time, or combine them with a simple vote. Pathologists don't work that way. They zoom out to read tissue structure and zoom in to read cells.

This project tests whether a model gets better when it does the same thing: 40X captures broad tissue architecture and organization, while 200X captures fine cellular morphology and nuclear detail.

Approach

Model comparison. Pre-trained ResNet50, EfficientNet, DenseNet, and Vision Transformer models were fine-tuned from ImageNet and compared against a custom CNN designed for histopathology textures, since tissue images differ a lot from natural photos.

Multi-scale fusion. Instead of voting, the fusion model combines magnification-specific CNN features late in the network, with attention that learns how much weight each magnification deserves.

Class imbalance. The data is 31.3% benign and 68.7% malignant. Focal loss and class weighting handled the imbalance, with stratified 70/15/15 train, validation, and test splits.

Results

ModelAccuracyPrecisionRecallF1
Fusion (40X + 200X)96.38%95.90%98.94%97.40%
ResNet50, 200X only94.70%96.17%96.17%96.17%
ResNet50, 40X only91.33%95.45%91.75%93.56%
Confusion matrices for 40X, 200X, and fusion models
Confusion matrices for each model. Fusion cuts missed malignant cases to 2.

Seeing inside the model

Grad-CAM highlights the regions that drove each prediction. On this malignant sample, the 40X model focused largely on empty background and predicted benign. At 200X, attention shifted onto the tissue and the prediction was correct.

Drag to compare the slide with the model's attention

Breast tissue slide, true label malignant Grad-CAM heatmap over the same slide SlideGrad-CAM
Actual: MalignantPredicted: Benign (51.9%)

Responsible AI check

The right metric. A missed cancer costs far more than a false alarm, so malignant recall was the priority metric, not overall accuracy. Fusion raised it to 98.94%.

Visible reasoning. Pathologists need to know why a model made a call before they can trust it. Grad-CAM makes each prediction inspectable, and it exposed exactly the kind of failure the fusion model was built to fix.

Stated limits. Trained and tested on one public dataset of 82 patients from a single institution. It has not been validated on slides from other labs or scanners, so it is a research model, not a diagnostic tool.