For independent identically distributed observations , the likelihood function and log-likelihood are
A maximum-likelihood estimator is a measurable choice , when a maximizer exists. It need not be unique. The one-observation score function is , and the Fisher information matrix is
Under the regularity assumptions, differentiation beneath the integral gives . Differentiating again gives the second information identity. Independent observations have total Fisher information .
Assume fixed parameter dimension, an interior true parameter , a positive-definite matrix , the standard differentiability and integrability conditions, and statistical consistency of . Then the asymptotic normality of a maximum likelihood estimator is
Here is the information per observation. Thus the leading covariance matrix of the estimator itself is .
To prove this, statistical consistency places the estimator in an interior ball about with probability tending to one, where its score function vanishes. Write and use the integral first-order Taylor formula for a vector map:
This integral matrix is needed in a vector problem; one does not have to assert a common scalar mean-value point for every score component. Put
The regular local uniform law of large numbers, continuity of the expected Hessian matrix, and statistical consistency give . The multivariate central limit theorem gives
The limiting information matrix is nonsingular, so . Solving the Taylor identity and applying the Slutsky theorem proves the displayed normal limit, since . Boundary parameters or singular information are outside this regular theorem.
Represent the subset model by a set of column indices and let contain those columns. Its ordinary least squares estimator sets the omitted coefficients to zero and has
For , use . Full column rank of implies full column rank for each . The matrix is an orthogonal projection matrix of rank .
Put and . Projection orthogonality gives the bias-variance decomposition of mean squared error
while the expected residual sum of squares is
These equations distinguish the mean-vector prediction risk in this problem from the additional noise in predicting a new response vector.
First take . Let project onto the full column space of . Since the full model contains , the residual estimate of Gaussian noise variance
is unbiased: its numerator has expectation . Consequently an unbiased Gaussian projection risk estimate is
Indeed its expectation is for every , even when the subset model omits true effects. Independence of the two terms is not needed for this expectation calculation. The full-model variance estimate is important: the subset residual variance generally includes omitted-variable bias.
For a finite candidate collection, choose a model minimizing , breaking ties in favour of a smaller model. The term is common to all models, so this amounts to minimizing
the Mallows Cp form. The fit improvement competes with an increasing dimension penalty. Unbiasedness holds for each fixed model; the estimate at the data-selected model need not remain unbiased, because selection favours downward fluctuations. The method is a heuristic model selection rule, rather than an assertion that it finds the true model with certainty.
The printed includes a genuine qualification. If , there are no full-model residual degrees of freedom. Except in the balanced case , no integrable data-only unbiased estimator can satisfy the requested identity for all means and unknown variances. Here is a proof of the unknown-variance risk estimation in a saturated Gaussian model obstruction. Since is invertible, ranges over all . Suppose had the required expectation at every and variance . For , let and let . Marginally . Applying the assumed identity conditionally gives
whereas applying it to the marginal distribution gives . Their difference is , which must vanish. Integrability under the marginal Gaussian justifies conditioning. If , the obstruction disappears and the residual sum of squares alone has expectation . Thus
The general construction therefore requires , or an independently available unbiased variance estimate.
The ordinary column Gram matrix is . Use the normalized empirical Gram matrix
so that standard Gaussian entries give . This is the normalization needed for concentration around .
The restricted isometry property of order with constant means that the normalized map approximately preserves the Euclidean norm of all vectors with at most nonzero coordinates:
Equivalently, every principal block with satisfies . The least such is its restricted isometry constant. In the unnormalized definition apply this property to itself.
For a standard normal variable , direct Gaussian integration gives for . Independence therefore yields the moment-generating function of a chi-squared distribution, centred here at its mean:
The inequality follows from . For , the Chernoff bound with gives
At the trivial probability bound suffices. This proves the requested bound, with a stronger prefactor one.
For the lower tail, gives for . Taking yields . Combining both tails gives the useful chi-squared concentration inequality
For set . A direct calculation shows
Thus the exponent is at least , proving
The same threshold bounds the two-sided tail.
Finally fix a deterministic and put . If , independence of the Gaussian rows gives independently. Therefore
The two-sided bound just proved supplies
For the quadratic form is deterministically zero; the strict inequality makes the formula valid in that case too. When , replacing the threshold by gives an absolute bound independent of . In particular every fixed such direction concentrates at rate for a fixed confidence level. The restricted isometry property requires a simultaneous statement over sparse directions; this fixed-direction calculation alone is not that stronger assertion.
Let . For any , the set
is compact. If it is empty there is nothing to prove. Otherwise continuity and the unique minimum give a strictly positive separation gap
Take any measurable attained minimizer of . Its defining inequality gives
Consequently
This proves argmin consistency under uniform convergence in probability. No continuity of is needed once the minimizer exists; compactness and continuity concern the deterministic separation gap. The uniform convergence in probability assumption controls all candidate parameters, including the random minimizer.
For the estimating equation, fix and put , . The prescribed signs make . Pointwise convergence in probability at just these two points implies
and similarly . With probability tending to one, . The intermediate value theorem then gives a zero inside , and uniqueness identifies it with . Hence
This is consistency of a uniquely bracketed zero. It needs no monotonicity, no continuity of the limit , and no uniform convergence of the . The bracket interval must lie in the domain of : read literally, the printed sign condition for every positive puts every real point into , so this requirement is satisfied. More generally it is enough to have an interval about and sign brackets arbitrarily close to it.

Articles by others on the same topic (0)

There are currently no matching articles.