Causal identification means that is uniquely determined by the observed joint distribution of under the causal assumptions. Equivalently, it admits an identifying formula containing only observed-data probabilities and conditional expectations.
The G-computation formula for this two-stage treatment isThe first factor is the observed mean outcome after the specified treatment history and intermediate value; the second averages over the intermediate-variable distribution generated after the first treatment.
In the second causal directed acyclic graph, depends on but not on , whereas depends on and the two latent roots are independent. HenceAlthough conditioning on conveys information about , the assignment uses only and fresh randomization, soThese are the two sequential exchangeability conditions. Using them successively, together with consistency of potential outcomes, giveswhich is the formula from part a.
Adding invalidates the argument in general. The latent variable confounds and , so need not equal the distribution of . When directly affects , that discrepancy no longer cancels after summing over . The same observed distribution can then correspond to different intervention means, so the displayed formula need not identify the effect.
Writefor the observed history just before . A sufficient condition is sequential exchangeabilityfor every treatment regime, together with consistency of potential outcomes and positivity in causal inference. Repeated conditioning then gives the longitudinal G-formulaGraphically, it is enough that each be D-separated from the final counterfactual under the specified regime after conditioning on its observed past. The two independences used in part b are precisely the instance.
- Instrument relevance: changes the conditional distribution of , for example on a set of positive probability.
- Instrumental-variable independence: is independent of latent outcome causes and the relevant potential outcomes, for example .
- The exclusion restriction: affects only through , written .
Together with consistency of potential outcomes and positivity in causal inference, these assumptions make variation in induced by causally interpretable. In the displayed graph, relevance is the edge , independence is the absence of a path from to after conditioning on , and exclusion is the absence of a direct edge.
Retaining the estimator exactly as printed, define the probability limitThe empirical equations are linear in . Their coefficient matrix converges toTherefore a sufficient condition, requiring neither parametric assumption, iswith finite moments sufficient for the weak law of large numbers. The empirical determinant then converges to a nonzero number, so the two linear equations have a unique solution with probability tending to one.
If Assumption 1 holds, put . Then , so conditional instrument independence givesConsequently both population estimating equations vanish at for any probability limit of .
If Assumption 2 holds, choose the linear-projection coefficientThen , whileThe first term is zero by conditional instrument independence and the second by the definition of . Thus solves the population equations. Under the nonsingularity condition from part b, the root is unique, so standard estimating equation consistency proveswhenever either Assumption 1 or Assumption 2 holds. This is double robustness.
LetAssumption 2 implies . Although the printed need not converge to , its first-order effect vanishes because . Inverting the Jacobian of the remaining estimating equations gives the influence functionThe first-stage condition yieldsHence the asymptotic variance of is the sandwich expression
There is a defect in the printed assumptions: alone does not determine , because depends on and can have a nonzero conditional mean given under unmeasured confounding. Under the standard intended strengthening , the formula simplifies toThe asymptotic variance of itself is .
A distribution is faithful to a Directed acyclic graph when every conditional independence in the distribution is implied by D-separation in . Together with the graphical Markov property, faithfulness makes conditional independence equivalent to D-separation and rules out independences caused only by exact parameter cancellation.
A backdoor path from to is a path whose first edge has an arrowhead at . The set satisfies the backdoor criterion when it contains no descendant of and blocks every backdoor path from to . Under this criterion, consistency, and positivity,so adjustment for identifies the intervention mean.
Definition 1 satisfies Property 1 but not Property 2. Let contain every non-descendant that blocks some backdoor path. Every backdoor path begins for a parent of ; whenever that path can transmit confounding, . Conditioning on therefore blocks every such path at its first nonendpoint vertex, so is sufficient.
For failure of Property 2, consider the faithful graph with arrowsThe variable blocks the backdoor path , so Definition 1 calls it a confounder. Every sufficient set containing must nevertheless contain to block . Once is included, deleting leaves a sufficient set. Thus can never be essential as Property 2 demands.
Definition 2 satisfies Property 2 but not Property 1. If belongs to every minimal sufficient adjustment set, choose one such set and put . By minimality, is sufficient and is not, proving Property 2.
For failure of Property 1, use the faithful chain-shaped backdoor pathBoth and are minimal sufficient adjustment sets. No variable belongs to every minimal sufficient set, so Definition 2 labels no variable a confounder, but the empty set is not sufficient.
Definition 3 satisfies Property 1 but not Property 2. Under faithfulness, its associational criterion contains enough non-descendants to block every open backdoor path. Indeed, if such a path remained open, its first unconditioned parent of would be D-connected to and, after a suitable conditioning set, to given ; faithfulness would place that parent in the Definition 3 set, a contradiction. Thus adjusting for all variables selected by Definition 3 is sufficient.
For failure of Property 2, considerwith both and observed and a faithful distribution. The instrumental variable is associated with . Conditioning on the collider opens , so is associated with given and Definition 3 calls it a confounder. Any sufficient set containing must also contain to block , but is already sufficient. Hence removing never destroys sufficiency, violating Property 2.
Restricting to pupils who already attended a Catholic middle school improves covariate balance: the treatment-group differences in family income, urban residence, and prior mathematics score are all much smaller. This makes severe extrapolation and measured confounding less prominent.
The restriction reduces the sample from to , so estimates are less precise. It also changes the target population to Catholic-middle-school pupils, reducing external validity for all United States pupils.
In the full sample, covariate adjustment moves the estimated coefficient from to , a large change consistent with substantial measured confounding. In the Catholic-middle-school subsample the estimates remain between and , supporting the claim that restriction has already improved comparability. The cost is visible in the standard errors, which rise from about -- to -- despite similar coefficient magnitudes.
A sufficient causal condition is conditional exchangeabilitytogether with consistency of potential outcomes and positivity in causal inference. For the ordinary-least-squares coefficient itself to equal one common causal effect, also require the correctly specified additive conditional-mean modelThen is the homogeneous treatment effect, and the adjusted coefficient consistently estimates both conditional effects and the average treatment effect in the analyzed population.
Let and let be a consistent estimate. The inverse-probability-weighted estimator of the average treatment effect isUnder the exchangeability, consistency, and positivity conditions in part ii, and consistent estimation of the propensity score, its probability limit isso it consistently estimates the average treatment effect. It does not require the additive outcome-regression model used to interpret the ordinary-least-squares coefficient.
In the bivariate probit model for endogenous treatment, affects treatment through its threshold equation and affects the outcome through its threshold equation. When , the two disturbances are dependent, so treatment status carries information about the latent outcome disturbance even after conditioning on . Consequentlyin general, and the no unmeasured confounding assumption fails.
Fix . For an observation with covariates , putand let be the standard-normal distribution function. The four conditional cell probabilities arewhere the first index is and the second is .
Define as any maximizer of the log likelihoodThen is the requested estimator for the fixed sensitivity value .
The correlation measures dependence between two normalized latent disturbances; it is not a scale-free measure of the strength of one physical confounder. Different latent-variable constructions can induce the same while producing different treatment-outcome confounding, and the same omitted cause can produce different after changing thresholds or disturbance scales. Moreover, re-estimating at each value of changes the entire latent model, so the fitted models do not represent one fixed data-generating mechanism with only its confounder strength varied. At the bivariate normal distribution is singular as well. Varying is a model-based sensitivity analysis, but interpreting the interval as an ordered range of strengths of a single unmeasured confounder is therefore logically unjustified.
Articles by others on the same topic
There are currently no matching articles.