For individual , let be the reported count, the gender indicator, and the minority indicator. The fitted Poisson regression isIts log-likelihood isThe maximum-likelihood estimator isHolding minority status fixed, changing the gender indicator from zero to one multiplies the fitted conditional expected value by . Thus the fitted mean count for men is about lower than that for women with the same minority status.
The researcher computed the Pearson chi-squared statisticand compared it with a distribution, using residual degrees of freedom. Under an adequate large-sample Poisson regression, should be roughly the residual degrees of freedom. The reported tail probability rounds numerically to zero and gives strong evidence of overdispersion.
Possible causes include unobserved heterogeneity or omitted covariates, dependence among respondents, excess zeros, or an incorrect mean function. The conclusion that the Poisson variance assumption fails is well supported, although this test alone does not identify the cause or establish that the Quasi-Poisson regression variance is correct.
With having rows and , the quasi-score equation isIt is the same coefficient equation as for Poisson maximum likelihood, explaining why the two models have identical coefficient estimates.
The Quasi-Poisson standard errors are trustworthy only if observations are independent, the log-linear mean is correct, and the variance is proportional to the mean with one common dispersion. Dependence, zero inflation, or covariate-dependent dispersion can invalidate this covariance formula.
A parametric bootstrap under model 1 proceeds as follows. Fit the Poisson model once and retain and the fitted means . For bootstrap repetition , independently drawrefit the same Poisson regression to , and save its gender estimate . The sample standard deviation of these estimates over many repetitions estimates the model-1 standard error. This bootstrap deliberately measures uncertainty under the fitted Poisson model; it does not repair real overdispersion unless the resampling model is enlarged to represent its cause.
The generalized linear mixed model tries to explain overdispersion by replacing the fixed minority coefficient with a Gaussian random intercept. Conditional on the group effect ,
This is a poor use of a random effect because minority has only two levels. Two realized intercepts contain almost no information about a random-effects distribution or its variance, and the two levels are substantively fixed categories rather than a sample from a population of groups. The model also has worse Akaike information criterion than model 1, , and its fit does not establish that the original overdispersion has disappeared.
Because models 1 and 3 are full likelihood models for the same response, one can compare their Akaike information criterion values; that favors model 1. One can also compare held-out count prediction by K-fold cross-validation, using a common loss such as Poisson deviance or negative log predictive density.
Articles by others on the same topic
There are currently no matching articles.