The problem
In 2022, monsoon floods in Pakistan displaced more than 30 million people and caused about $30 billion in damage, with roughly ten times normal rainfall.
This project asks whether classic time series algorithms, trained on more than a century of data, can recognize when rainfall is moving far outside its normal range.

Approach
Five models were compared across three datasets: monthly rainfall from 1901 to 2016, World Bank climate indicators, and NOAA's Oceanic Niño Index for El Niño and La Niña effects.
ARIMA(2,0,2) served as the non-seasonal baseline. SARIMA(1,0,2)(1,1,1,12) added the 12-month monsoon cycle. SARIMAX added ocean temperature data. Two naive baselines, the historical average and a repeat of last year, kept the models honest.
Results
| Model | RMSE (mm) | MAE (mm) |
|---|---|---|
| SARIMA(1,0,2)(1,1,1,12) | 14.06 | 10.98 |
| SARIMAX (+ El Niño data) | 13.70 | 10.62 |
| Historical average | 14.09 | 11.01 |
| ARIMA(2,0,2) | 16.64 | 13.48 |
| Last year repeat | 27.63 | 18.38 |

The warning signal
For July 2022 the model predicted about 57 mm of rain. The actual total was roughly ten times that. The forecast wasn't supposed to predict the flood. The point is that rainfall this far outside the confidence interval is the early warning.


Responsible AI check
Honest baselines. The naive historical average came within 0.03 mm RMSE of SARIMA. I report that openly: the model’s value is in its seasonal structure and uncertainty band, not a dramatic accuracy gain.
Extra data isn’t automatically better. El Niño data confirmed a real relationship with rainfall but barely improved the forecast, so the simpler model stays the headline.
Stated limits. National monthly totals miss river flow, local terrain, and daily intensity. The warning signal is a promising hypothesis, not a deployed system.