Skip to content
Back to selected work

Case study / 2025

Breast Cancer Prediction Web App

A Streamlit classifier that labels a breast mass malignant or benign from 30 cell-nucleus measurements, tuned so that the errors it makes are the cheaper kind.

Primary result

97.07% malignant recall at a cost-weighted threshold, against a baseline that catches none

  • Python
  • Scikit-learn
  • SVM (RBF)
  • Streamlit
  • Pandas
  • NumPy
Breast Cancer Prediction Web App2025 / shipped
BENIGN CLUSTERMALIGNANT CLUSTER30 ATTRIBUTES · 569 SAMPLES
97.07%
Malignant recall
Malignant recall
0.9957
ROC-AUC
ROC-AUC
62.74%
Baseline accuracy
Baseline accuracy
0.225
Tuned threshold
Tuned threshold

Challenge

The problem to solve

Accuracy is the wrong headline for this dataset. It is 62.7% benign, so a model that answers "benign" every time scores 62.74% and catches no cancers at all. The real target is recall on malignancy — and a decision threshold chosen to reflect what a miss actually costs.

Approach

Technical direction

Compared logistic regression, an RBF-kernel SVM and a random forest against an explicit majority-class baseline, scored with repeated stratified cross-validation — 10 folds by 5 repeats — rather than a single split. Selected the SVM on malignant recall with ROC-AUC as the tiebreaker, then moved the decision threshold off 0.5 by weighting a false negative as ten times costlier than a false alarm.

Outcomes

Key outcomes

  1. 01Scored every model against an explicit majority-class baseline — 62.74% ± 0.70 accuracy, 0% malignant recall, 0.500 ROC-AUC — so the headline number has something to beat.
  2. 02Selected an RBF-kernel SVM at 97.54% ± 1.72 accuracy, 97.07% ± 3.27 malignant recall and 0.9957 ROC-AUC, ahead of logistic regression (97.44%) and random forest (96.35%).
  3. 03Used repeated stratified cross-validation, 10 folds × 5 repeats, so each figure carries a spread rather than resting on one lucky seed.
  4. 04Tuned the decision threshold to 0.225 by weighting false negatives 10× false positives: nine more false alarms in exchange for two more cancers caught.
  5. 05Fixed a leakage bug where the scaler was fitted before the train-test split, and consolidated separate model and scaler pickles into a single pipeline artifact.

Process

Process & architecture

01

Why accuracy is the wrong headline

The Wisconsin diagnostic dataset is 62.7% benign. A model that answers "benign" unconditionally therefore scores 62.74% accuracy while catching zero cancers. Any accuracy figure quoted without that number beside it is close to meaningless, which is why the baseline is reported alongside every model here.

02

The data

569 samples, each described by 30 numeric features computed from digitised images of cell nuclei — radius, texture, perimeter, area, smoothness and so on, each as a mean, a standard error and a worst-case value.

03

Scoring that survives a reshuffle

A single train-test split on 569 rows moves by percentage points depending on the seed. Every model here is scored with repeated stratified cross-validation — 10 folds, 5 repeats — so each result comes with a standard deviation and the comparison between models means something.

04

Model comparison

Against the 62.74% baseline: logistic regression reached 97.44% ± 1.93 accuracy with 96.68% malignant recall; the RBF-kernel SVM reached 97.54% ± 1.72 with 97.07% ± 3.27 recall and 0.9957 ROC-AUC; the random forest trailed at 96.35% with 93.88% recall. The SVM was selected on recall, with ROC-AUC breaking the tie against logistic regression.

05

Moving the threshold off 0.5

The default 0.5 cutoff silently assumes a false alarm and a missed cancer cost the same. They do not. Weighting a false negative as ten times costlier puts the optimal threshold at 0.225 — which on this test set trades nine additional false alarms for two additional malignancies detected. That is the trade a screening tool should be making.

06

A leakage bug worth naming

An earlier version fitted the feature scaler before splitting into train and test, letting test-set statistics inform the transform and flattering every score that followed. The scaler now lives inside the pipeline so it is refitted within each fold, and the model and scaler ship as one artifact rather than two pickles that could drift apart.

07

Application interface

A Streamlit front end takes the clinical measurements and returns a malignant/benign call with its probability, using the tuned threshold rather than the default so the interface reflects the same cost assumption as the evaluation.

08

What it is not

This is a classifier trained on one public dataset of 569 cases from a single source, evaluated by cross-validation rather than against an external cohort. It demonstrates a decision process. It is not a diagnostic instrument.