Observed heterogeneity is variation in outcome propensities explained by measured covariates, such as age or sex. In logistic regression, subjects with different recorded predictor values may have different success probabilities, even before allowing for their previous outcomes. This is heterogeneity visible through observed predictors.
Unobserved heterogeneity is persistent variation between subjects arising from unmeasured characteristics. A subject-specific latent variable or random intercept can represent it. Subjects with high latent success propensities tend to succeed repeatedly, so their observed outcomes can have positive serial correlation even when outcomes are conditionally independent given that propensity. Apparent persistence therefore need not imply true state dependence.
True state dependence, also called true contagion, means that a previous outcome changes the distribution of a subsequent outcome after controlling both measured covariates and persistent unobserved heterogeneity. For example, an earlier successful week may make later success more likely through habit formation. A lagged-outcome coefficient in an inadequate statistical model can also reflect omitted subject differences, so positive observed persistence alone does not distinguish true state dependence from unobserved heterogeneity.
Let indicate a smoking-free week, the recorded sex indicator, age in years and assigned treatment. The screening history supplies ; earlier lag initializations are also zero. A history-dependent logistic regression for the first model is
Here , is the observed past, and the five unknown regression coefficients have time-invariant values. Conditional on baseline covariates, this model has the first-order Markov property: only the immediately preceding outcome enters the current conditional probability. Distinct subjects have independent histories. There is no subject random effect or additional time trend in this fit.
Using the chain rule for probabilities, the individual conditional likelihood is
The lagged values in are the individual's actual preceding outcomes. This product is a sequential conditional likelihood, not an assertion of unconditional independence of the ten readings. The full conditional likelihood is ; no extra Bernoulli distribution factor is attached to the fixed screening history.
The second logistic regression adds while retaining the first lag and baseline covariates. The models are nested under . Their likelihood-ratio test statistic is the reduction in binomial deviance,
Under the null and regular large-sample conditions for the correctly specified conditional likelihood, this has an approximate chi-squared distribution with one statistical degree of freedom. Since , reject at 5%; the approximate p-value is . Prefer the two-lag model to the one-lag model. Its additional lag captures statistically useful information in the history.
The cumulative-response logistic model uses instead of the two individual lag predictors. It and the two-lag history-dependent logistic regression are nonnested: the entire accumulated history generally cannot be represented using only the two recent outcomes. Their binomial deviance difference therefore has no ordinary nested chi-squared distribution calibration.
Use the Akaike information criterion, which up to the same saturated-model constant equals for these Bernoulli distribution conditional likelihoods. The two-lag fit has six coefficients, while the cumulative fit has five:
Thus prefer the cumulative-history model, whose Akaike information criterion is lower by 35 despite its smaller number of statistical parameters. This is a model-selection comparison of conditional histories, not a nested likelihood-ratio test.
The preferred cumulative-response logistic model estimates
Its logit link describes conditional smoking-free probability given baseline predictors and prior successful weeks. Holding the other predictors fixed, males have times the female success odds; the reported p-value is , and the approximate 95% odds ratio confidence interval is . Age has estimated odds ratio per additional year, with p-value , giving little evidence for an age association in this fit.
The combined treatment has conditional success odds ratio versus the reference treatment, with approximate 95% confidence interval and p-value . The fitted conditional odds are about 39% higher for the combined treatment. The trial randomization supports treatment comparisons, but conditioning on accumulated post-treatment outcomes means this coefficient is not directly the marginal total treatment effect.
Each previous successful week multiplies current success odds by , with approximate 95% confidence interval and very small p-value . This is strong fitted persistence. It can reflect true state dependence, unobserved heterogeneity, or an omitted calendar-time trend; this fit alone cannot distinguish them. The baseline intercept implies success probability for a reference-treatment female aged zero with no previous success. That age is outside the study's useful interpretation range, so the intercept chiefly anchors the regression. The Bernoulli distribution dispersion parameter is fixed at one, and the residual binomial deviance is 1122 on 995 statistical degrees of freedom. With individual binary outcomes, comparing that binomial deviance mechanically to a chi-squared distribution is not a reliable general goodness-of-fit test.
Take the event to mean three smoking-free weeks in succession, . The initial cumulative count is zero. Along this path it equals in weeks one, two, three, respectively. In the cumulative-response logistic model, define
The chain rule for probabilities then gives
The three conditional probabilities are approximately . If “stop during the first three weeks” instead means at least one smoking-free week by week three, its different event has probability : along the all-failure path the cumulative count stays zero. Stating the event resolves this wording ambiguity.

Articles by others on the same topic (0)

There are currently no matching articles.