Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2014/iii/paper-30/2/solution

Represent the subset model by a set of column indices and let contain those columns. Its ordinary least squares estimator sets the omitted coefficients to zero and has
For , use . Full column rank of implies full column rank for each . The matrix is an orthogonal projection matrix of rank .
Put and . Projection orthogonality gives the bias-variance decomposition of mean squared error
while the expected residual sum of squares is
These equations distinguish the mean-vector prediction risk in this problem from the additional noise in predicting a new response vector.
First take . Let project onto the full column space of . Since the full model contains , the residual estimate of Gaussian noise variance
is unbiased: its numerator has expectation . Consequently an unbiased Gaussian projection risk estimate is
Indeed its expectation is for every , even when the subset model omits true effects. Independence of the two terms is not needed for this expectation calculation. The full-model variance estimate is important: the subset residual variance generally includes omitted-variable bias.
For a finite candidate collection, choose a model minimizing , breaking ties in favour of a smaller model. The term is common to all models, so this amounts to minimizing
the Mallows Cp form. The fit improvement competes with an increasing dimension penalty. Unbiasedness holds for each fixed model; the estimate at the data-selected model need not remain unbiased, because selection favours downward fluctuations. The method is a heuristic model selection rule, rather than an assertion that it finds the true model with certainty.
The printed includes a genuine qualification. If , there are no full-model residual degrees of freedom. Except in the balanced case , no integrable data-only unbiased estimator can satisfy the requested identity for all means and unknown variances. Here is a proof of the unknown-variance risk estimation in a saturated Gaussian model obstruction. Since is invertible, ranges over all . Suppose had the required expectation at every and variance . For , let and let . Marginally . Applying the assumed identity conditionally gives
whereas applying it to the marginal distribution gives . Their difference is , which must vanish. Integrability under the marginal Gaussian justifies conditioning. If , the obstruction disappears and the residual sum of squares alone has expectation . Thus
The general construction therefore requires , or an independently available unbiased variance estimate.

New to topics? Read the docs here!