Case study / 2025
Fire Risk Classification from Weather
A Flask service that estimates the probability of a fire day in northern Algeria from six weather inputs, trained only on variables that are observed rather than derived from the label.
Primary result
0.938 ROC-AUC from weather alone, against a 0.500 baseline
- Python
- Scikit-learn
- Random Forest
- GroupKFold
- Flask
- Joblib
- 0.938ROC-AUC
- ROC-AUC
- 88.1%Accuracy
- Accuracy
- 89.8%Recall
- Recall
- 243Days of data
- Days of data
Challenge
The problem to solve
The obvious way to model this dataset scores extremely well and means nothing. The Fire Weather Index is a formula computed from the same weather already in the table, and the published fire labels are close to a single threshold on it. Any model handed those columns is reading the answer back. The problem was to build something that still works once the shortcut is removed.
Approach
Technical direction
Restricted the features to temperature, relative humidity, wind speed, rainfall, month and region, dropping all six FWI-system outputs. Compared logistic regression against a random forest under GroupKFold over region-month, so an entire region-month is held out at once and correlated consecutive days cannot straddle a fold. The selected model is served from a Flask app with an HTML form, a JSON endpoint and a health check.
Outcomes
Key outcomes
- 01Found two leakage paths in the original framing: FWI is reproducible from ISI and BUI at R² = 0.985, and the single rule FWI > 3.5 reproduces the published fire labels 94.2% of the time.
- 02Rebuilt the task as fire-occurrence classification from six observed weather variables, excluding FFMC, DMC, DC, ISI, BUI and FWI.
- 03Validated with GroupKFold over region-month instead of random splits, so adjacent days from the same region and month cannot appear on both sides of a fold.
- 04Random forest reached 88.1% accuracy, 89.8% recall and 0.938 ROC-AUC against an always-fire baseline of 56.4% accuracy and 0.500 ROC-AUC.
- 05Served the model through Flask with a form, a JSON /api/predict endpoint returning a fire probability, and a /health check.
Process
Process & architecture
What the data is
Two regions of northern Algeria, June to September 2012, 243 days in total, each labelled fire or not fire. Fire days make up 56.4% of the record, so a model that answers "fire" every time is already right 56.4% of the time. That is the number worth beating.
The shortcut that had to go
The dataset ships with the Fire Weather Index system — FFMC, DMC, DC, ISI, BUI and FWI. None of these are observations; they are computed from the same weather columns. Regressing FWI on ISI and BUI returns R² = 0.985. The label is compromised too: the single rule FWI > 3.5 reproduces the published fire labels 94.2% of the time.
What honesty costs
Trained with the FWI columns included, the model reaches 94.2% accuracy and 0.992 ROC-AUC — the same figure the FWI > 3.5 rule reaches on its own, without any model at all. That accuracy is the label being handed back, so this project gives it up and reports the harder number instead.
Features that survive
Six inputs remain: temperature, relative humidity, wind speed, rainfall, month and region. Each is either measured directly or known in advance, so a prediction made from them is a prediction rather than a restatement.
Validation that respects place and time
Fire weather persists for days, so a random split can put Tuesday in training and Wednesday in test and flatter the result. GroupKFold over region-month holds out a whole region-month at once. The gap is visible in both models: the optimistic random split scores the random forest at 0.943 ROC-AUC and logistic regression at 0.890, against 0.938 and 0.877 under grouping.
Model comparison
Logistic regression reached 78.6% accuracy and 0.877 ROC-AUC. The random forest reached 88.1% accuracy, 89.8% recall and 0.938 ROC-AUC, and was selected — recall carries the weight here, since a missed fire day costs more than a false alarm.
Serving
A Flask app loads the joblib artifact produced by train.py and exposes three surfaces: an HTML form, POST /api/predict returning a fire probability and a verdict at the 0.5 threshold, and GET /health.
Limits worth stating
One fire season, two regions, 243 rows. The model estimates fire occurrence on the same day as the weather it is given, so it is a classifier and not a forecast. The provenance of the labels is undocumented, which is exactly why the FWI threshold agreement matters.