Population Stability Index (PSI)
Population Stability Index (PSI) is a statistical metric used in machine learning and credit risk modeling to quantify data drift and population shift between two distributions—typically a reference baseline dataset (training/validation) and a target dataset (production/serving).
Formula
Given a continuous feature or model score divided into buckets/bins:
Per-Bin PSI:
Where:
- Expected% (): The percentage of observations in bin in the baseline/reference population (e.g., training data).
- Actual% (): The percentage of observations in bin in the production/target population.
Total Population PSI:
Note: PSI is symmetric and mathematically equivalent to the symmetrized Kullback-Leibler (KL) divergence: .
Interpretation Thresholds
| PSI Score | Drift Severity | Recommended Action |
|---|---|---|
| No significant shift | Population has not changed meaningfully; continue monitoring. | |
| Moderate shift | Minor drift detected; investigate root cause and prepare for retraining. | |
| Significant shift | Major population drift; model predictions may be degraded. Trigger model retraining or recalibration. |
Python Calculation Example
import numpy as np
def calculate_psi(expected, actual, num_buckets=10, epsilon=1e-4):
"""
Calculate Population Stability Index between expected (baseline) and actual (target) samples.
"""
# Create equal-frequency bins from the baseline distribution
percentiles = np.linspace(0, 100, num_buckets + 1)
bins = np.percentile(expected, percentiles)
bins[0] = -np.inf
bins[-1] = np.inf
# Calculate proportions in each bin
expected_counts, _ = np.histogram(expected, bins=bins)
actual_counts, _ = np.histogram(actual, bins=bins)
expected_pct = expected_counts / len(expected)
actual_pct = actual_counts / len(actual)
# Handle zero division / log(0) with small epsilon
expected_pct = np.clip(expected_pct, epsilon, 1.0)
actual_pct = np.clip(actual_pct, epsilon, 1.0)
# PSI formula
psi_values = (actual_pct - expected_pct) * np.log(actual_pct / expected_pct)
return np.sum(psi_values)
Best Practices
- Use 10–20 Quantile Bins: Compute bucket thresholds from the baseline population so that each baseline bin holds approximately equal sample sizes ().
- Handle Zero Counts with Smoothing: If a bin has zero observations in actual or expected distributions, is undefined. Use an (e.g., ) smoothing floor.
- Monitor Both Features & Output Scores: Track PSI on output probability predictions (prediction drift) as well as individual critical features (covariate shift).