The useful autoregressive model is for the squares, not for signed observations. Put and
Then
The errors form a martingale difference sequence relative to the noise history, because the current standardized noise is independent of the past. When the fourth moment is finite they have finite variance and are uncorrelated across distinct times, although their conditional variance depends on the regressor. In centered form, .
For ordinary least squares, use the response-regressor pairs , , . With their separate means and , minimize . If the regressor sum of squares is positive, the estimators are
The two means use the matched pairs; replacing them indiscriminately by a single full-sample mean is not the exact least-squares formula. The noise-series representation supplies an ergodic stationary process. If , finite regressor second moments and the error's zero conditional expectation justify the usual population regression and statistical consistency argument. The observations still define a finite-sample least-squares fit outside that moment range, but the ordinary finite-variance justification must not be claimed there. If parameter constraints are required, minimize the same criterion subject to and , rather than assert that unconstrained estimates automatically satisfy them.
First, the absolute values in the printed limit are an error: the correct long-run variance of a stationary process is the signed sum of autocovariances. Directly,
so
Each coefficient tends to and has absolute value at most . Absolute summability and the dominated convergence theorem therefore give
The sum of absolute values is an upper bound, not the general limit. For a concrete counterexample to the printed assertion, take with unit-variance iid noise. Then , , and other covariances vanish. The absolute sum is , but , so .
For correlated data the central limit theorem can still hold under suitable strong mixing of a stationary process and moment conditions, but the limiting variance is the long-run variance of a stationary process, rather than the one-observation variance. When it is positive,
Positive serial dependence usually increases the standard error, while negative dependence can reduce it. The approximate effective sample size of a stationary sample is when that denominator is positive. A long-memory time series can require a different normalization or a different limit law. Absolute covariance summability alone does not prove a CLT: if with a common independent random scale taking values and with equal probabilities and iid with the standard normal distribution, the off-diagonal covariances are zero, yet has the nonnormal scale-mixture law . The example is not an ergodic stationary process. If the long-run variance of a stationary process is zero, the usual nondegenerate square-root- CLT is unavailable.
For the causal AR(1), and
The conditional distribution follows because the current innovation is independent of the past.
For a fixed known , condition on the observed and use the transitions . Their conditional maximum likelihood criterion is, up to constants,
Differentiating in yields
Since ,
Thus it is an unbiased estimator, even conditionally on , and
It has statistical consistency with mean-square convergence, and the iid-noise strong law of large numbers also gives almost-sure statistical consistency. Its conditional distribution is exactly normal with the displayed mean and variance. If is also observed and all transitions are used, replace by .
The fixed- qualification is necessary for the exact finite-sample claims. If is jointly estimated, conditional likelihood function is linear regression with an intercept: writing and , the unconstrained estimators are
This ratio is not generally an unbiased estimator and does not have the preceding finite-sample variance. Under the usual stationary regression conditions it has statistical consistency; its asymptotic variance is . Indeed, it differs from by , an asymptotically negligible endpoint term. Profiling an unknown innovation variance does not change the fixed- estimate of .