For a model with maximized log likelihood , fitted parameters, and observations,
Both combine goodness of fit with a complexity penalty; smaller values are preferred.
Both criteria are used for model selection. AIC estimates relative out-of-sample predictive or Kullback-Leibler risk and is attractive when prediction is the main aim. BIC approximates a log Bayes factor under regular fixed-dimensional models and is consistent for selecting a true finite-dimensional model when one is present. Their penalties differ by versus . For , BIC penalizes each additional parameter more strongly and therefore tends to select smaller models.
For model , let contain its active columns. With known noise variance one,
and full column rank gives
Let be orthogonal projection onto the column space of , and let be the true mean. Then
Up to a common constant,
The true model has and ; every other candidate has and . Its excess expected AIC is , while its excess expected BIC is when . Thus either minimum expected criterion selects the true model.
With one true regressor and one additional candidate regressor, the reduction in residual sum of squares from fitting the larger nested model is . AIC chooses the wrong larger model exactly when
whose probability is fixed and positive, independent of . BIC chooses it when
For this probability is strictly smaller than the AIC error probability, and it tends to zero as .

Articles by others on the same topic (0)

There are currently no matching articles.