Skip to content
Back to selected work

Case study / 2026

Customer Churn Prediction & Retention Dashboard

A churn classifier over 7,043 telecom accounts that reports what it costs to catch churners rather than hiding behind an accuracy figure, served through a Streamlit scoring app.

Primary result

80.4% of churners caught, at roughly one real save per two calls

  • Python
  • Scikit-learn
  • Pandas
  • Streamlit
  • Joblib
Customer Churn Prediction & Retention Dashboard2026 / shipped
80.4% CAUGHT7,043 CUSTOMERS · 0.845 ROC-AUCLOW RETENTION SIGNALHIGH CHURN SIGNAL
0.845
ROC-AUC
ROC-AUC
80.4%
Churners caught
Churners caught
51.5%
Flag precision
Flag precision
7,043
Customers
Customers

Challenge

The problem to solve

Accuracy is the wrong target here. 73.5% of these customers do not churn, so a model that answers "will not churn" every time scores 73.46% and identifies nobody worth calling. The useful question is how many real churners you can catch, and how many pointless retention calls that costs.

Approach

Technical direction

Compared logistic regression and a random forest against an explicit never-churn baseline, both with balanced class weights, scored under repeated stratified cross-validation at 5 splits × 3 repeats. All preprocessing lives inside the pipeline so the scaler and encoder refit within each fold. Logistic regression ships — not because it wins on headline numbers, but for seven points more recall and coefficients a retention team can read.

Outcomes

Key outcomes

  1. 01Scored every model against a never-churn baseline — 73.46% accuracy, 0% recall, 0.500 ROC-AUC — so the headline figures have something to beat.
  2. 02Logistic regression reached 0.8449 ± 0.0107 ROC-AUC, 80.35% ± 2.16 recall on churners and 74.71% accuracy under repeated stratified cross-validation.
  3. 03Chose it over the random forest despite the forest scoring marginally higher on ROC-AUC (0.8459, inside the spread), because it gives up seven points of recall on the class that costs money.
  4. 04Reported the operating cost plainly: 1,501 churners caught against 1,422 false alarms, i.e. about one genuine save per two retention calls.
  5. 05Established that MonthlyCharges cannot be read from the coefficients, because TotalCharges is tenure × MonthlyCharges at R² = 0.9991 and the collinearity splits one signal across three columns.

Process

Process & architecture

01

Why not accuracy

The dataset is 73.5% non-churners. A model that predicts "stays" for everyone scores 73.46% accuracy, beats a naive expectation, and is completely worthless — it flags nobody. That baseline is measured and reported alongside every model here, because an accuracy number without it says nothing.

02

The data

The public Telco Customer Churn dataset: 7,043 customers, 19 features covering contract, tenure, services and charges, with a binary churn label. The base churn rate is 26.5%.

03

Cleaning, and one decision worth naming

TotalCharges arrives as text and is blank for 11 customers — every one of them at tenure zero, meaning they have never been billed. They are kept with TotalCharges set to 0 rather than dropped, because a brand-new customer is precisely the case a retention model has to handle.

04

Scoring that survives a reshuffle

A single train-test split on this data moves by more than a percentage point depending on the seed. Every figure comes from RepeatedStratifiedKFold at 5 splits × 3 repeats, so each carries a standard deviation, and the preprocessing is refitted inside each fold rather than on the full dataset.

05

Choosing the model that loses on paper

The random forest scores 0.8459 ROC-AUC to logistic regression's 0.8449 — a gap of 0.001 against a spread of ±0.011, which is noise. What is not noise is recall: 73.07% for the forest against 80.35% for logistic regression. Since the whole point is catching churners, the regression ships, and its coefficients double as the retention story.

06

What the errors cost

Out-of-fold across all 7,043 customers: 1,501 churners caught, 368 missed, and 1,422 customers flagged who were never going to leave. That is 51.5% precision — roughly one real save per two retention calls. Whether that trade is worth making depends on the cost of a call against the value of a retained customer, which is a business decision. The model surfaces it rather than burying it.

07

Drivers, and one that does not survive scrutiny

Tenure dominates as a protective factor (−1.34), followed by two-year contracts (−0.81); month-to-month contracts (+0.62) and fibre-optic internet (+0.52) push the other way. MonthlyCharges lands at −0.42, which does not mean higher bills reduce churn — TotalCharges is tenure × MonthlyCharges at R² = 0.9991, and under collinearity that tight the individual coefficients become unstable and flip sign. Tenure and contract type are the pair worth trusting.

08

What it is not

This dataset is a snapshot with no time dimension, so the model classifies customers who have already churned rather than forecasting who will churn next month. No temporal validation is possible, and calling it a forecast would be overclaiming.