logo
CHEATSHEETSTOOLSABOUT中文

Loss Functions & Regularization in Regression

In statistical modeling and machine learning, regression algorithms combine a loss function (measuring prediction discrepancy) with a regularization penalty (constraining model complexity).

1. Regression Loss Functions

Loss Function Mathematical Form L(y,y^)\mathcal{L}(y, \hat{y}) Robust to Outliers? Target Statistic Common Algorithm
Squared Loss (L2L_2) 12(y−y^)2\frac{1}{2}(y - \hat{y})^2 No (quadratic penalty on large errors) Mean Ordinary Least Squares, Ridge
Absolute Loss (L1L_1) ∥y−y^∥\|y - \hat{y}\| Yes (linear penalty on large errors) Median Quantile Regression, Robust Regression
Huber Loss {12(y−y^)2if ∥y−y^∥≤δδ∥y−y^∥−12δ2otherwise\begin{cases} \frac{1}{2}(y - \hat{y})^2 & \text{if } \|y - \hat{y}\| \le \delta \\ \delta \|y - \hat{y}\| - \frac{1}{2}\delta^2 & \text{otherwise} \end{cases} Yes Hybrid Huber Regression
ϵ\epsilon-Insensitive Loss max⁡(0,∥y−y^∥−ϵ)\max(0, \|y - \hat{y}\| - \epsilon) Yes Support vectors Support Vector Regression (SVR)
Log Loss (Cross-Entropy) ln⁡(1+e−y⋅y^)\ln(1 + e^{-y \cdot \hat{y}}) Moderate Probability Logistic Regression
Hinge Loss max⁡(0,1−y⋅y^)\max(0, 1 - y \cdot \hat{y}) Moderate Margin separator Support Vector Machine (SVM)

2. Regularization Penalties

Regularization constraints prevent overfitting and stabilize parameter estimation in the presence of collinear features:

min⁡θ∑i=1mL(y(i),hθ(x(i)))+R(θ)\min_{\theta} \sum_{i=1}^m \mathcal{L}(y^{(i)}, h_\theta(x^{(i)})) + \mathcal{R}(\theta)

L2L_2 Regularization (Ridge / Weight Decay)

R(θ)=λ2∥θ∥22=λ2∑j=1nθj2\mathcal{R}(\theta) = \frac{\lambda}{2} \|\theta\|_2^2 = \frac{\lambda}{2} \sum_{j=1}^n \theta_j^2
  • Geometric Constraint: Spherical (L2L_2 ball).
  • Behavior: Shrinks coefficients smoothly towards zero without setting them strictly to zero.
  • Property: Has a closed-form analytical solution when used with squared loss: θ=(XTX+λI)−1XTy\theta = (X^T X + \lambda I)^{-1} X^T y.

L1L_1 Regularization (Lasso)

R(θ)=λ∥θ∥1=λ∑j=1n∣θj∣\mathcal{R}(\theta) = \lambda \|\theta\|_1 = \lambda \sum_{j=1}^n |\theta_j|
  • Geometric Constraint: Polyhedral / Diamond-shaped (L1L_1 diamond).
  • Behavior: Due to sharp corners on axes, contour intersections frequently land exactly on coordinate planes, driving uninformative feature weights to strictly zero.
  • Property: Acts as an automatic feature selection mechanism, producing sparse models.

Elastic Net Regularization

Combines L1L_1 and L2L_2 penalties via mixing parameter α∈[0,1]\alpha \in [0, 1]:

R(θ)=λ(α∥θ∥1+1−α2∥θ∥22)\mathcal{R}(\theta) = \lambda \left( \alpha \|\theta\|_1 + \frac{1 - \alpha}{2} \|\theta\|_2^2 \right)
  • Solves Lasso Limitations: When multiple features are strongly correlated, Lasso tends to arbitrarily pick one and discard the rest. Elastic Net retains the entire group of correlated features while maintaining sparsity.