Series 1 wanders over a changing level rather than fluctuating around a stable local mean. Its sample autocorrelation function is strongly positive and decreases very slowly. This is the usual diagnostic evidence for an ordinary unit root: an autoregressive polynomial containing , with a zero at , and a stationary model after first differencing. The plots support an integrated model, rather than specifying the number of its remaining stationary autoregressive or moving-average terms.
Series 2 has a pronounced oscillation with period about six observations. Its sample autocorrelation alternates between large positive and negative values with little damping: approximately positive at multiples of six and negative halfway between. Together with the changing amplitude, this suggests a conjugate pair of unit-circle zeros near
The associated real autoregressive factor is . A targeted filter removes this pair; the broader seasonal difference operator also contains it but introduces additional differencing factors. This is the oscillatory unit-root diagnosis from an undamped sample autocorrelation.
Thus Series 1 suggests a zero at 1; Series 2 suggests a conjugate pair on the unit circle at a seasonal frequency. These are model diagnoses, not deductions of exact roots from a finite sample. A stationary model very close to a unit root can look similar, and an undamped periodic covariance can also arise from a stationary random sinusoid. The figure does not identify exact orders or prove nonstationarity by itself.
Among the supplied candidates, choose ARMA(2,1). It has the smallest Akaike information criterion. Relative to ARMA(2,2), its AIC improvement is ; the additional second moving-average estimate is only half a standard error from zero and increases the log likelihood function by only about . Dropping it is a reasonable parsimony choice, although that AIC gap is small. ARMA(1,1) has AIC larger by , a much clearer loss of fit. In the chosen model the second autoregressive term is about seven standard errors from zero, so it should not be dropped merely to obtain order one. The first autoregressive term is less precisely estimated; this does not justify automatically deleting it without fitting and comparing the reduced candidate.
Using the usual positive-sign moving-average convention, the fitted autoregressive moving-average model is
These are plug-in values rounded as given, not exact population parameters. The model has zero mean. Its polynomials are and .
With angular frequency , so that the autocovariance is , the spectral density of a stationary process is
The frequency convention makes the normalization unambiguous.
The autoregressive polynomial factors as , with zeros
Both have modulus greater than one. The causality root criterion for an autoregressive model therefore gives a causal stationary solution. The moving-average polynomial has its only zero at , also outside the unit circle, so the invertibility of a moving-average model holds. There is no common root to cancel.
To use unit-variance white noise, put . One suitable pair is
Then with . The factor two changes the innovation scale, not the zero of the moving-average polynomial.
Write , with innovation variance four. The generating function of this linear process is
Equating coefficients gives , and for . Thus the first five coefficients, including lag zero, are
For all remaining lags, partial fractions give
The coefficients are absolutely summable, so the series converges in mean square and defines the causal moving-average expansion. In the unit-variance convention of part (c), with ; its first five coefficients are . This is the two-geometric-coefficient expansion of a causal ARMA(2,1) process.
White noise orthogonality gives, for any integer ,
To see this directly, expand the covariance of the two convergent series. Only matching noise indices contribute. Cauchy-Schwarz inequality makes the coefficient-product sum finite, justifying the covariance limit.
There is also a closed expression. Set , , , . For ,
and negative lags follow by symmetry. Dividing by the expression at zero gives the same autocorrelation function.
Interpret stationarity in the usual second-order time-series sense and assume nondegenerate noise, . For a two-sided autoregressive equation, the missing existence condition is
It is important to separate this from causality. If , the unique stationary solution is . If , there is still a stationary solution, but it is anticausal:
Both expansions converge in L2 because their coefficients are square summable. Substitution verifies the equation. Their means are zero and their covariance functions depend only on lag. Uniqueness follows by iterating the equation backward in the first case and forward in the second: the remainders or tend to zero in L2 for any stationary finite-variance solution. This is the stationary versus causal solution of a two-sided AR(1) equation.
For , iteration gives
The variance of the right side is . The variance of the left side is at most by stationarity and Cauchy-Schwarz inequality. These are incompatible as . Thus no weakly stationary finite-variance solution exists at those unit roots.
If the intended claim includes a causal innovation representation, its condition is instead , as in the next part. The stated white noise equation alone does not say that is orthogonal to the past of . If zero innovation variance is allowed, the unit-root exclusion has degenerate exceptions, such as random constant solutions when ; the nondegenerate convention is necessary for the asserted nonexistence.
The condition for a causal linear-filter solution is . Its mean-square expansion is . Summing the matching white noise terms gives
The assumed orthogonality of every to every extends to every by L2 convergence of that expansion. Hence the added-noise process has mean zero and
This depends only on lag, proving weak stationarity. Its autocorrelation has the same geometric tail as the latent autoregression, but its positive-lag correlations are reduced by the additional variance at lag zero. This is the autocovariance of an AR(1) process observed with white noise. Strict stationarity or Gaussianity is not implied by white noise covariance assumptions alone.
Apply to the observed process and call the result . Then
Its only nonzero covariance lags are
Seek an invertible moving-average factor with . Matching these covariances requires and . Solving gives
These are the three requested parameters in the positive-sign moving-average convention. For nonzero total noise variance,
so and the larger quadratic root gives the invertible factor.
Covariance matching alone would not identify arbitrary processes in distribution. To obtain an actual representation on the given space, define
The series converges in L2. The spectrum of is , so this filtered process has constant spectrum and is white noise. Thus
This is the invertible ARMA factorization of an AR(1)-plus-noise process. If , it reduces to AR(1); if , it reduces to white noise. If and , then , the common factor cancels and . Therefore the orders are at most (1,1); no unnecessarily minimal-order claim is made in those degenerate cases.
Take the latent autoregression as the scalar state . A state-space model is
The transition and observation matrices are both scalar, , . State-noise variance is , observation-noise variance is , and the cross-noise covariance is zero at every pair of times.
A complete stationary initialization is
It is orthogonal to future state noise and to all observation noise. This is the stationary initialization of a scalar linear state-space model. This specifies the initial state in terms of the actual given two-sided white noise sequence, as well as its second-order law; simply starting from zero would give transient rather than stationary observations.
If a Gaussian state-space specification is intended, the complete specialization is , independent of the future iid Gaussian state and observation noises, themselves independent with variances . Under the printed assumptions alone, Gaussian distributions and independence cannot be deduced from white noise orthogonality; the equations and stationary-series initialization above give the exact second-order representation without adding them.
Generate independent and use the Box-Muller transform
For , , so its density is on . The angle is uniform on and independent of . Dividing the joint radial-angular density by the polar-coordinate Jacobian gives
Thus both outputs are independent and have the standard normal distribution. Independent pairs of uniforms give further independent normal observations. The null event is excluded in implementation so the logarithm is finite.
For fixed parameters , every new error is independent of the previous observations. Put . Since , the first three conditional densities are
and
In general,
Its density is . The recursion is interpreted from , as required to define from the two given initial values. These are parameter-conditional sampling distributions; integrating over the prior would instead give predictive mixtures.
Use the chain rule for conditional densities. With , the likelihood function of the nondegenerate observations is
The first residual is and carries no parameter information. There is no ordinary joint Lebesgue density for because deterministically; this expression is the conditional likelihood function given the fixed initial values, or the likelihood function on . The initial point mass is parameter independent and has no effect on inference. This is the conditional likelihood of an initialized Gaussian AR(2) process.
Define the sufficient data sums, all over ,
Multiplying the likelihood function by the independent standard-normal priors and collecting the quadratic terms gives
Let and . This precision matrix is , where , so it is positive definite even for a short or singular design. Completing the square proves
The fully normalized posterior density is
This is Gaussian conjugacy for an initialized AR(2) regression.
Completing each one-dimensional square, or using the supplied conditional-normal identity, gives
When , all regressor sums vanish and the posterior remains the independent standard-normal prior. No stationary-parameter restriction is imposed: the specified prior is on all of , and the finite initialized chain is defined for all .
Start from any . For independent draws from the standard normal distribution at each sweep, the Gibbs sampler can be implemented as
The second update must use the newly drawn first coordinate. The transform in part (a) supplies the needed independent Gaussian draws. Each conditional update preserves the joint posterior, so their composition does too.
Convergence is particularly transparent here. Centering at the posterior mean gives
Positive definiteness gives , so this is a stable Gaussian autoregression. The full sampler has the desired posterior as its limiting invariant law. This is the linear contraction of a two-coordinate Gaussian Gibbs sweep.
For a posterior-integrable function , its posterior expectation is estimated after a burn-in by
The ergodic theorem justifies this Monte Carlo average. The retained values of also approximate its posterior distribution and quantiles. If a Bayesian point estimate is requested under squared-error loss, the posterior mean is the appropriate estimate, when its needed moments exist. An arbitrary function need not have a finite posterior mean; integrability must be assumed for the displayed target. Successive Gibbs draws are correlated, so uncertainty in the Monte Carlo average should use chain-aware error estimates rather than treating them as iid.
Let where , and set it to zero on the g-null set where . The domination assumption makes there and gives . Integrating it also gives . For iid proposals , the importance sampling estimator is
Cauchy-Schwarz inequality under gives . Direct integration proves unbiasedness, . Moreover
Thus
The iid central limit theorem applies to these finite-variance weighted observations:
If the variance is zero, this denotes the point mass at zero. This is the bounded-weight importance-sampling moment bound.
For each proposal , independently draw and accept it when
The ratio is at most one, as required. For any measurable set ,
so the acceptance probability is and the conditional distribution of an accepted proposal has density . Repeating independent trials until acceptance therefore gives an exact draw from , and repeating the whole procedure gives iid target draws. This proves rejection sampling.
The number of proposals for one successful draw is geometric with mean . Thus is also the expected proposal cost of this exact simulation method.
Apply independent uniforms to the proposals, and let indicate acceptance. Each is Bernoulli with success probability , and the indicators are independent because the pairs are independent. Hence
For a fixed acceptance pattern, the values at accepted positions are independent with density , by the single-trial conditional density calculation. The same product law holds for every pattern of a given size. Consequently, conditional on , the ordered accepted observations have joint density . This is the binomial count and iid values in fixed-budget rejection sampling; the count provides no information about those target values.
There is a finite-sample empty-output event, with probability . Any estimator dividing by must be defined separately on .
Write and . The Bernoulli central limit theorem gives
If , all proposals are accepted and this limit is degenerate.
Conditional on , the accepted sample is iid from . Since , its size tends to infinity, and the ordinary target-sample CLT therefore gives
Assign any fixed value to the estimator on ; the probability of that event tends to zero, so it does not change this limit. This is the random-count central limit theorem for accepted rejection samples.
For comparison at the same budget of proposals, Slutsky's theorem rescales this as
The distinction between accepted-sample size and proposal count is essential for the final efficiency comparison.
At a common proposal budget, the two asymptotic variances are
Thus prefer the estimator with the smaller proposal-budget variance; the stated assumptions do not give a universal winner. Importance sampling uses every proposal and avoids an empty accepted sample, but those facts alone do not establish variance dominance. The bound from part (a) only says . If , it does imply that importance sampling is at least as efficient asymptotically.
For explicit nonconstant counterexamples in both directions, take and on , and zero outside, with the sharp envelope . For , direct integration gives , and
The accepted-sample mean is better. For the same densities but , the target variance is unchanged, while and
Importance sampling is now better. This is the variance reversal under additive shifts of an importance-sampling integrand.
There is a useful distinction about the Rao-Blackwell theorem. The weighted importance estimator is the conditional expectation, given the proposals, of the fixed-denominator rejection estimator . It is not the conditional expectation of the random-denominator mean . Rao-Blackwell variance reduction for the former therefore does not prove a comparison with the latter. This is the Rao-Blackwell identity for a fixed-denominator rejection estimator. The appropriate answer is the proposal-budget variance comparison of importance and rejection sampling, with the empty-output convention specified.

Articles by others on the same topic (0)

There are currently no matching articles.