logo

Population Stability Index (PSI)

Population Stability Index (PSI) is a statistical metric used in machine learning and credit risk modeling to quantify data drift and population shift between two distributions—typically a reference baseline dataset (training/validation) and a target dataset (production/serving).

Formula

Given a continuous feature or model score divided into kk buckets/bins:

Per-Bin PSI:

PSIi=(Actual%i−Expected%i)×ln⁡(Actual%iExpected%i)\text{PSI}_i = \big( \text{Actual}\%_i - \text{Expected}\%_i \big) \times \ln \left( \frac{\text{Actual}\%_i}{\text{Expected}\%_i} \right)

Where:

  • Expected% (EiE_i): The percentage of observations in bin ii in the baseline/reference population (e.g., training data).
  • Actual% (AiA_i): The percentage of observations in bin ii in the production/target population.

Total Population PSI:

PSI=∑i=1kPSIi=∑i=1k(Ai−Ei)×ln⁡(AiEi)\text{PSI} = \sum_{i=1}^{k} \text{PSI}_i = \sum_{i=1}^{k} \big( A_i - E_i \big) \times \ln \left( \frac{A_i}{E_i} \right)

Note: PSI is symmetric and mathematically equivalent to the symmetrized Kullback-Leibler (KL) divergence: DKL(A∥E)+DKL(E∥A)D_{\text{KL}}(A \parallel E) + D_{\text{KL}}(E \parallel A).

Interpretation Thresholds

PSI Score Drift Severity Recommended Action
PSI<0.10\text{PSI} < 0.10 No significant shift Population has not changed meaningfully; continue monitoring.
0.10≤PSI<0.250.10 \le \text{PSI} < 0.25 Moderate shift Minor drift detected; investigate root cause and prepare for retraining.
PSI≥0.25\text{PSI} \ge 0.25 Significant shift Major population drift; model predictions may be degraded. Trigger model retraining or recalibration.

Python Calculation Example

import numpy as np

def calculate_psi(expected, actual, num_buckets=10, epsilon=1e-4):
    """
    Calculate Population Stability Index between expected (baseline) and actual (target) samples.
    """
    # Create equal-frequency bins from the baseline distribution
    percentiles = np.linspace(0, 100, num_buckets + 1)
    bins = np.percentile(expected, percentiles)
    bins[0] = -np.inf
    bins[-1] = np.inf

    # Calculate proportions in each bin
    expected_counts, _ = np.histogram(expected, bins=bins)
    actual_counts, _ = np.histogram(actual, bins=bins)

    expected_pct = expected_counts / len(expected)
    actual_pct = actual_counts / len(actual)

    # Handle zero division / log(0) with small epsilon
    expected_pct = np.clip(expected_pct, epsilon, 1.0)
    actual_pct = np.clip(actual_pct, epsilon, 1.0)

    # PSI formula
    psi_values = (actual_pct - expected_pct) * np.log(actual_pct / expected_pct)
    return np.sum(psi_values)

Best Practices

  1. Use 10–20 Quantile Bins: Compute bucket thresholds from the baseline population so that each baseline bin holds approximately equal sample sizes (5–10%5\text{–}10\%).
  2. Handle Zero Counts with Smoothing: If a bin has zero observations in actual or expected distributions, ln⁡(0)\ln(0) is undefined. Use an ϵ\epsilon (e.g., 10−410^{-4}) smoothing floor.
  3. Monitor Both Features & Output Scores: Track PSI on output probability predictions (prediction drift) as well as individual critical features (covariate shift).