Let denote the full data and the pattern of indicators, with for an observed component and for a missing one. For each fixed pattern , split the data as . The missing at random condition is
Here parameterizes the missingness mechanism. Equivalently its conditional probability, with the observed data fixed, is constant over all possible completions of the missing data. The restriction is pattern-specific because the observed components depend on . Missing at random allows missingness to depend on observed values; missing completely at random imposes independence from all the data.
Ask the doctor whether, among patients with the same first-year drug-use status, patients who dropped out would be more or less likely to have used the drug in the second year than those who remained. The comparison is within the observed first-year categories, not just between all dropouts and all completers.
If there is no remaining relationship with second-year drug use after conditioning on first-year use, missing at random is plausible for these recorded variables. If treatment-related difficulties, new medication or deterioration lead to dropout in a way not accounted for by the recorded first-year status, missing at random may fail. A different response rate in the two first-year categories is allowed under missing at random. The unobserved second-year outcomes mean that the observed table alone cannot establish this assumption; clinical knowledge or additional follow-up information is needed.
Apply MAR standardization over a fully observed covariate. All first-year statuses are known, so the estimated probabilities of no use and use in year one are and . Within those categories, missing at random lets the observed second-year probabilities represent the corresponding dropout outcomes as well. They are estimated by and .
The law of total probability then gives
Equivalently, impute expected drug-use counts and for the two dropout groups, add these to the 25 observed second-year users, and divide by 102. This uses both the complete records and the fully observed first-year information.
Under missing completely at random, the observed second-year outcomes form a random subsample, so the direct pooled complete-case analysis estimate is
It is valid under missing completely at random because dropout no longer changes the marginal distribution of second-year use. Pooling is simpler than reporting separate conditional probabilities, but MCAR alone does not make this estimate more efficient than the estimate in the previous part. The fully observed first-year variable carries useful outcome information and can improve precision even when dropout is completely random.
To make the qualification explicit, write , , and for the constant response probability. For patients the leading variance of the pooled estimator is . The standardization estimator in the preceding part has leading variance
Indeed its first-order centered contribution is : the two terms have zero covariance, and their variances give that expression. The law of total variance therefore yields
The inequality is strict when there is dropout and first-year use predicts second-year use, as the distinct conditional rates suggest here. Thus the requested universal efficiency claim needs qualification. In the unrestricted joint binary-outcome model, maximizing the observed-data likelihood under either missing at random or missing completely at random gives the same standardization estimate : the all-patient first-year proportion and the two observed conditional second-year proportions maximize its factored outcome likelihood. Restricting the distinct missingness parameters to a common response probability affects their factor, not this estimate. The pooled value is the simple valid MCAR answer; retaining the first-year information gives the efficient MCAR answer and does not require replacing the previous estimate.
An ignorable missingness mechanism need not be absent from the data-generating process. It means that likelihood inference about the data-model parameters can omit the missingness factor, while still integrating over missing values.
Write the complete joint data statistical probability density as and the conditional missingness probability as . All values and only are observed. Under missing at random, is constant as varies with fixed. The actual observed-data likelihood is therefore
The assumed distinctness is understood as independent variation of and . Maximizing over multiplies by a factor independent of ; likelihood ratios, scores and likelihood curvature for are therefore unchanged by omitting . This proves likelihood ignorability.
For implementation of the observed-data likelihood with a missing covariate, a linear regression model for must be accompanied by an appropriate model for the distribution of . For independent individuals, write for the regression statistical probability density and for the age statistical probability density. Up to the ignorable factor,
Ignorability does not authorize discarding missing-age cases or assuming their ages have the distribution seen in the complete cases. It removes the need to model the observation mechanism for likelihood inference under the stated conditions, not the need to handle the missing covariates.

Articles by others on the same topic (0)

There are currently no matching articles.