Loss Functions & Regularization in Regression
In statistical modeling and machine learning, regression algorithms combine a loss function (measuring prediction discrepancy) with a regularization penalty (constraining model complexity).
1. Regression Loss Functions
| Loss Function | Mathematical Form | Robust to Outliers? | Target Statistic | Common Algorithm |
|---|---|---|---|---|
| Squared Loss () | No (quadratic penalty on large errors) | Mean | Ordinary Least Squares, Ridge | |
| Absolute Loss () | Yes (linear penalty on large errors) | Median | Quantile Regression, Robust Regression | |
| Huber Loss | Yes | Hybrid | Huber Regression | |
| -Insensitive Loss | Yes | Support vectors | Support Vector Regression (SVR) | |
| Log Loss (Cross-Entropy) | Moderate | Probability | Logistic Regression | |
| Hinge Loss | Moderate | Margin separator | Support Vector Machine (SVM) |
2. Regularization Penalties
Regularization constraints prevent overfitting and stabilize parameter estimation in the presence of collinear features:
Regularization (Ridge / Weight Decay)
- Geometric Constraint: Spherical ( ball).
- Behavior: Shrinks coefficients smoothly towards zero without setting them strictly to zero.
- Property: Has a closed-form analytical solution when used with squared loss: .
Regularization (Lasso)
- Geometric Constraint: Polyhedral / Diamond-shaped ( diamond).
- Behavior: Due to sharp corners on axes, contour intersections frequently land exactly on coordinate planes, driving uninformative feature weights to strictly zero.
- Property: Acts as an automatic feature selection mechanism, producing sparse models.
Elastic Net Regularization
Combines and penalties via mixing parameter :
- Solves Lasso Limitations: When multiple features are strongly correlated, Lasso tends to arbitrarily pick one and discard the rest. Elastic Net retains the entire group of correlated features while maintaining sparsity.