Keep baseline covariates fixed and put , , and for . Integration of a Poisson count over the Gamma component gives
This negative binomial distribution has mean and variance . Therefore the unconditional one-term distribution is the zero-inflated negative binomial distribution
In particular its zero probability is , not simply . As this becomes a Zero-inflated Poisson distribution.
The mixing effect has , and . The law of total variance and law of total covariance now give
Equivalently, for , and . This also supplies raw second moments by adding . Dependence arises from both shared heterogeneity and shared structural zeros, even when the positive random effect is constant.
The printed moment variance matches the hierarchical variance at the same parameter values, but its printed off-diagonal covariance is . Equality would require , or . It is therefore false to claim that the two approaches consistently estimate all the same parameters in general.
The precise distinction is between consistency of mean ratios and identification of structural parameters. In the hierarchical model, with , both the quadratic variance coefficient and the covariance coefficient in terms of the marginal mean are
For the moment parameterization, writing its parameters as , these coefficients are and . Its intercept is identifiable from the mean only as . Matching the true first two moments demands . One solution always is
For there is also a solution , , , with the corresponding shifted intercept. Indeed eliminating gives . This is moment aliasing in a shared zero-inflated count model: matching the moments does not identify the actual structural-zero probability.
For a concrete counterexample take . Then and . The only admissible moment-matching choice is , with an intercept shifted by ; it cannot recover the true structural-zero proportion 0.2. Correct marginal means can give consistent term and baseline-covariate slopes with sandwich standard errors; they do not justify consistent estimation of the intended , or latent intercept. Nor does a mean-only estimating equation identify at all. Full-distribution inference provides information beyond these first two moments.