Solution (source code)

= Solution

For <independent> identically distributed observations $Y_1,\ldots,Y_n$, the <likelihood function> and <log-likelihood> are
$$
L_n(\theta)=\prod_{i=1}^n f(Y_i,\theta),\qquad\ell_n(\theta)=\sum_{i=1}^n\log f(Y_i,\theta).
$$
A <maximum-likelihood estimator> is a measurable choice $\widehat\theta_n\in\operatorname{argmax}_{\theta\in\Theta}L_n(\theta)$, when a maximizer exists. It need not be unique. The one-observation <score function> is $s_\theta(y)=\nabla_\theta\log f(y,\theta)$, and the <Fisher information matrix> is
$$
I(\theta)=\mathbb E_\theta[s_\theta(Y)s_\theta(Y)^T]=-\mathbb E_\theta[\nabla_\theta^2\log f(Y,\theta)].
$$
Under the regularity assumptions, differentiation beneath the integral gives $\mathbb E_\theta s_\theta=\nabla_\theta\int f=0$. Differentiating again gives the second information identity. <Independent> observations have total <Fisher information> $nI(\theta)$.

Assume fixed parameter dimension, an interior true parameter $\theta_0$, a <positive-definite matrix> $I(\theta_0)$, the standard differentiability and integrability conditions, and <statistical consistency> of $\widehat\theta_n$. Then the <asymptotic normality of a maximum likelihood estimator> is
$$
\boxed{\sqrt n(\widehat\theta_n-\theta_0)\ \xrightarrow{d}\ N_p(0,I(\theta_0)^{-1}).}
$$
Here $I$ is the information per observation. Thus the leading <covariance matrix> of the estimator itself is $I(\theta_0)^{-1}/n$.

To prove this, <statistical consistency> places the estimator in an interior ball about $\theta_0$ with <probability> tending to one, where its <score function> vanishes. Write $d_n=\widehat\theta_n-\theta_0$ and use the <integral first-order Taylor formula for a vector map>:
$$
0=\nabla\ell_n(\theta_0)+\left\{\int_0^1\nabla^2\ell_n(\theta_0+t d_n)\,dt\right\}d_n.
$$
This integral matrix is needed in a vector problem; one does not have to assert a common scalar mean-value point for every score component. Put
$$
J_n=-\frac1n\int_0^1\nabla^2\ell_n(\theta_0+t d_n)\,dt.
$$
The regular local <uniform law of large numbers>, <continuity> of the expected <Hessian matrix>, and <statistical consistency> give $J_n\xrightarrow{P}I(\theta_0)$. The <multivariate central limit theorem> gives
$$
\frac1{\sqrt n}\nabla\ell_n(\theta_0)=\frac1{\sqrt n}\sum_{i=1}^n s_{\theta_0}(Y_i)\ \xrightarrow{d}\ N_p(0,I(\theta_0)).
$$
The limiting information matrix is nonsingular, so $J_n^{-1}\xrightarrow{P}I(\theta_0)^{-1}$. Solving the Taylor identity and applying the <Slutsky theorem> proves the displayed normal limit, since $I^{-1}II^{-1}=I^{-1}$. Boundary parameters or singular information are outside this regular theorem.