Model selection compares candidate statistical models and chooses one according to a criterion balancing fit, predictive performance and complexity.
For a full-rank normal linear model with error variance , the maximum-likelihood residual variance satisfiesConsequently
For a normal linear model with known error variance , the scaled Akaike criterion differs by a model-independent constant from
If are independent vectors and is a rank- orthogonal projection, thenThe first term is squared approximation bias, while is fitted-model variance and is irreducible new-response noise.
If is the rank- orthogonal projection onto a normal linear model's column space, thenAdding gives , so Mallows' is unbiased for independent-copy prediction error even when the projection model is misspecified.
For a model with fitted parameters, observations and maximized likelihood , the Bayesian information criterion is
Articles by others on the same topic
Model selection is the process of choosing the most appropriate statistical or machine learning model for a specific dataset and task. The objective is to identify a model that best captures the underlying patterns in the data while avoiding overfitting or underfitting. This process is crucial because different models can yield different predictions and insights from the same data.