For a model with fitted parameters and maximized likelihood , the Akaike information criterion is
Backward selection starts with the full model, deletes the single variable giving the smallest AIC when that AIC is lower than the current value, and repeats until no deletion improves it. The first deletion is "indus", which gives AIC .
Adding variables can reduce training bias and increase the maximized likelihood, but it also increases estimation variance and optimism. The penalty estimates this optimism, so minimizing AIC implements a bias-variance tradeoff aimed at expected out-of-sample Kullback–Leibler performance.
For ,
At the maximum-likelihood estimates,
so
There are regression coefficients and one variance parameter. Adding twice this parameter count gives
The common terms cancel, and
This is negative exactly when
There are three regression coefficients and residual degrees of freedom, so . The reported residual standard error uses the unbiased divisor:
Hence the maximum-likelihood variance estimate is . Substituting this, , and the four fitted parameters—three coefficients plus —into the formula in part c(ii) gives the AIC.

Articles by others on the same topic (0)

There are currently no matching articles.