logo

Population Stability Index (PSI)

Population Stability Index (PSI) is a statistical metric used in machine learning and credit risk modeling to quantify data drift and population shift between two distributions—typically a reference baseline dataset (training/validation) and a target dataset (production/serving).

Formula

Given a continuous feature or model score divided into k k buckets/bins:

Per-Bin PSI:

PSI i = ( Actual % i − Expected % i ) × ln ⁡ ( Actual % i Expected % i ) \text{PSI}_i = \big( \text{Actual}\%_i - \text{Expected}\%_i \big) \times \ln \left( \frac{\text{Actual}\%_i}{\text{Expected}\%_i} \right)

Where:

  • Expected% ( E i E_i ): The percentage of observations in bin i i in the baseline/reference population (e.g., training data).
  • Actual% ( A i A_i ): The percentage of observations in bin i i in the production/target population.

Total Population PSI:

PSI = ∑ i = 1 k PSI i = ∑ i = 1 k ( A i − E i ) × ln ⁡ ( A i E i ) \text{PSI} = \sum_{i=1}^{k} \text{PSI}_i = \sum_{i=1}^{k} \big( A_i - E_i \big) \times \ln \left( \frac{A_i}{E_i} \right)

Note: PSI is symmetric and mathematically equivalent to the symmetrized Kullback-Leibler (KL) divergence: D KL ( A ∥ E ) + D KL ( E ∥ A ) D_{\text{KL}}(A \parallel E) + D_{\text{KL}}(E \parallel A) .

Interpretation Thresholds

PSI Score Drift Severity Recommended Action
PSI < 0.10 \text{PSI} < 0.10 No significant shift Population has not changed meaningfully; continue monitoring.
0.10 ≤ PSI < 0.25 0.10 \le \text{PSI} < 0.25 Moderate shift Minor drift detected; investigate root cause and prepare for retraining.
PSI ≥ 0.25 \text{PSI} \ge 0.25 Significant shift Major population drift; model predictions may be degraded. Trigger model retraining or recalibration.

Python Calculation Example

import numpy as np

def calculate_psi(expected, actual, num_buckets=10, epsilon=1e-4):
    """
    Calculate Population Stability Index between expected (baseline) and actual (target) samples.
    """
    # Create equal-frequency bins from the baseline distribution
    percentiles = np.linspace(0, 100, num_buckets + 1)
    bins = np.percentile(expected, percentiles)
    bins[0] = -np.inf
    bins[-1] = np.inf

    # Calculate proportions in each bin
    expected_counts, _ = np.histogram(expected, bins=bins)
    actual_counts, _ = np.histogram(actual, bins=bins)

    expected_pct = expected_counts / len(expected)
    actual_pct = actual_counts / len(actual)

    # Handle zero division / log(0) with small epsilon
    expected_pct = np.clip(expected_pct, epsilon, 1.0)
    actual_pct = np.clip(actual_pct, epsilon, 1.0)

    # PSI formula
    psi_values = (actual_pct - expected_pct) * np.log(actual_pct / expected_pct)
    return np.sum(psi_values)

Best Practices

  1. Use 10–20 Quantile Bins: Compute bucket thresholds from the baseline population so that each baseline bin holds approximately equal sample sizes ( 5 – 10 % 5\text{–}10\% ).
  2. Handle Zero Counts with Smoothing: If a bin has zero observations in actual or expected distributions, ln ⁡ ( 0 ) \ln(0) is undefined. Use an ϵ \epsilon (e.g., 10 − 4 10^{-4} ) smoothing floor.
  3. Monitor Both Features & Output Scores: Track PSI on output probability predictions (prediction drift) as well as individual critical features (covariate shift).