Case study / 2026
Customer Churn Prediction & Retention Dashboard
A churn classifier over 7,043 telecom accounts that reports what it costs to catch churners rather than hiding behind an accuracy figure, served through a Streamlit scoring app.
Primary result
80.4% of churners caught, at roughly one real save per two calls
- Python
- Scikit-learn
- Pandas
- Streamlit
- Joblib
- 0.845ROC-AUC
- ROC-AUC
- 80.4%Churners caught
- Churners caught
- 51.5%Flag precision
- Flag precision
- 7,043Customers
- Customers
Challenge
The problem to solve
Accuracy is the wrong target here. 73.5% of these customers do not churn, so a model that answers "will not churn" every time scores 73.46% and identifies nobody worth calling. The useful question is how many real churners you can catch, and how many pointless retention calls that costs.
Approach
Technical direction
Compared logistic regression and a random forest against an explicit never-churn baseline, both with balanced class weights, scored under repeated stratified cross-validation at 5 splits × 3 repeats. All preprocessing lives inside the pipeline so the scaler and encoder refit within each fold. Logistic regression ships — not because it wins on headline numbers, but for seven points more recall and coefficients a retention team can read.
Outcomes
Key outcomes
- 01Scored every model against a never-churn baseline — 73.46% accuracy, 0% recall, 0.500 ROC-AUC — so the headline figures have something to beat.
- 02Logistic regression reached 0.8449 ± 0.0107 ROC-AUC, 80.35% ± 2.16 recall on churners and 74.71% accuracy under repeated stratified cross-validation.
- 03Chose it over the random forest despite the forest scoring marginally higher on ROC-AUC (0.8459, inside the spread), because it gives up seven points of recall on the class that costs money.
- 04Reported the operating cost plainly: 1,501 churners caught against 1,422 false alarms, i.e. about one genuine save per two retention calls.
- 05Established that MonthlyCharges cannot be read from the coefficients, because TotalCharges is tenure × MonthlyCharges at R² = 0.9991 and the collinearity splits one signal across three columns.
Process
Process & architecture
Why not accuracy
The dataset is 73.5% non-churners. A model that predicts "stays" for everyone scores 73.46% accuracy, beats a naive expectation, and is completely worthless — it flags nobody. That baseline is measured and reported alongside every model here, because an accuracy number without it says nothing.
The data
The public Telco Customer Churn dataset: 7,043 customers, 19 features covering contract, tenure, services and charges, with a binary churn label. The base churn rate is 26.5%.
Cleaning, and one decision worth naming
TotalCharges arrives as text and is blank for 11 customers — every one of them at tenure zero, meaning they have never been billed. They are kept with TotalCharges set to 0 rather than dropped, because a brand-new customer is precisely the case a retention model has to handle.
Scoring that survives a reshuffle
A single train-test split on this data moves by more than a percentage point depending on the seed. Every figure comes from RepeatedStratifiedKFold at 5 splits × 3 repeats, so each carries a standard deviation, and the preprocessing is refitted inside each fold rather than on the full dataset.
Choosing the model that loses on paper
The random forest scores 0.8459 ROC-AUC to logistic regression's 0.8449 — a gap of 0.001 against a spread of ±0.011, which is noise. What is not noise is recall: 73.07% for the forest against 80.35% for logistic regression. Since the whole point is catching churners, the regression ships, and its coefficients double as the retention story.
What the errors cost
Out-of-fold across all 7,043 customers: 1,501 churners caught, 368 missed, and 1,422 customers flagged who were never going to leave. That is 51.5% precision — roughly one real save per two retention calls. Whether that trade is worth making depends on the cost of a call against the value of a retained customer, which is a business decision. The model surfaces it rather than burying it.
Drivers, and one that does not survive scrutiny
Tenure dominates as a protective factor (−1.34), followed by two-year contracts (−0.81); month-to-month contracts (+0.62) and fibre-optic internet (+0.52) push the other way. MonthlyCharges lands at −0.42, which does not mean higher bills reduce churn — TotalCharges is tenure × MonthlyCharges at R² = 0.9991, and under collinearity that tight the individual coefficients become unstable and flip sign. Tenure and contract type are the pair worth trusting.
What it is not
This dataset is a snapshot with no time dimension, so the model classifies customers who have already churned rather than forecasting who will churn next month. No temporal validation is possible, and calling it a forecast would be overclaiming.