Causal identification means that is uniquely determined by the observed joint distribution of under the causal assumptions. Equivalently, it admits an identifying formula containing only observed-data probabilities and conditional expectations.
The G-computation formula for this two-stage treatment is
The first factor is the observed mean outcome after the specified treatment history and intermediate value; the second averages over the intermediate-variable distribution generated after the first treatment.
In the second causal directed acyclic graph, depends on but not on , whereas depends on and the two latent roots are independent. Hence
Although conditioning on conveys information about , the assignment uses only and fresh randomization, so
These are the two sequential exchangeability conditions. Using them successively, together with consistency of potential outcomes, gives
which is the formula from part a.
Adding invalidates the argument in general. The latent variable confounds and , so need not equal the distribution of . When directly affects , that discrepancy no longer cancels after summing over . The same observed distribution can then correspond to different intervention means, so the displayed formula need not identify the effect.
Write
for the observed history just before . A sufficient condition is sequential exchangeability
for every treatment regime, together with consistency of potential outcomes and positivity in causal inference. Repeated conditioning then gives the longitudinal G-formula
Graphically, it is enough that each be D-separated from the final counterfactual under the specified regime after conditioning on its observed past. The two independences used in part b are precisely the instance.
Conditionally on , a valid instrumental variable must satisfy three core conditions.
Together with consistency of potential outcomes and positivity in causal inference, these assumptions make variation in induced by causally interpretable. In the displayed graph, relevance is the edge , independence is the absence of a path from to after conditioning on , and exclusion is the absence of a direct edge.
Retaining the estimator exactly as printed, define the probability limit
The empirical equations are linear in . Their coefficient matrix converges to
Therefore a sufficient condition, requiring neither parametric assumption, is
with finite moments sufficient for the weak law of large numbers. The empirical determinant then converges to a nonzero number, so the two linear equations have a unique solution with probability tending to one.
The homogeneous treatment effect and consistency imply
Instrument validity gives .
If Assumption 1 holds, put . Then , so conditional instrument independence gives
Consequently both population estimating equations vanish at for any probability limit of .
If Assumption 2 holds, choose the linear-projection coefficient
Then , while
The first term is zero by conditional instrument independence and the second by the definition of . Thus solves the population equations. Under the nonsingularity condition from part b, the root is unique, so standard estimating equation consistency proves
whenever either Assumption 1 or Assumption 2 holds. This is double robustness.
Let
Assumption 2 implies . Although the printed need not converge to , its first-order effect vanishes because . Inverting the Jacobian of the remaining estimating equations gives the influence function
The first-stage condition yields
Hence the asymptotic variance of is the sandwich expression
There is a defect in the printed assumptions: alone does not determine , because depends on and can have a nonzero conditional mean given under unmeasured confounding. Under the standard intended strengthening , the formula simplifies to
The asymptotic variance of itself is .
A distribution is faithful to a Directed acyclic graph when every conditional independence in the distribution is implied by D-separation in . Together with the graphical Markov property, faithfulness makes conditional independence equivalent to D-separation and rules out independences caused only by exact parameter cancellation.
A backdoor path from to is a path whose first edge has an arrowhead at . The set satisfies the backdoor criterion when it contains no descendant of and blocks every backdoor path from to . Under this criterion, consistency, and positivity,
so adjustment for identifies the intervention mean.
Definition 1 satisfies Property 1 but not Property 2. Let contain every non-descendant that blocks some backdoor path. Every backdoor path begins for a parent of ; whenever that path can transmit confounding, . Conditioning on therefore blocks every such path at its first nonendpoint vertex, so is sufficient.
For failure of Property 2, consider the faithful graph with arrows
The variable blocks the backdoor path , so Definition 1 calls it a confounder. Every sufficient set containing must nevertheless contain to block . Once is included, deleting leaves a sufficient set. Thus can never be essential as Property 2 demands.
Definition 2 satisfies Property 2 but not Property 1. If belongs to every minimal sufficient adjustment set, choose one such set and put . By minimality, is sufficient and is not, proving Property 2.
For failure of Property 1, use the faithful chain-shaped backdoor path
Both and are minimal sufficient adjustment sets. No variable belongs to every minimal sufficient set, so Definition 2 labels no variable a confounder, but the empty set is not sufficient.
Definition 3 satisfies Property 1 but not Property 2. Under faithfulness, its associational criterion contains enough non-descendants to block every open backdoor path. Indeed, if such a path remained open, its first unconditioned parent of would be D-connected to and, after a suitable conditioning set, to given ; faithfulness would place that parent in the Definition 3 set, a contradiction. Thus adjusting for all variables selected by Definition 3 is sufficient.
For failure of Property 2, consider
with both and observed and a faithful distribution. The instrumental variable is associated with . Conditioning on the collider opens , so is associated with given and Definition 3 calls it a confounder. Any sufficient set containing must also contain to block , but is already sufficient. Hence removing never destroys sufficiency, violating Property 2.
Restricting to pupils who already attended a Catholic middle school improves covariate balance: the treatment-group differences in family income, urban residence, and prior mathematics score are all much smaller. This makes severe extrapolation and measured confounding less prominent.
The restriction reduces the sample from to , so estimates are less precise. It also changes the target population to Catholic-middle-school pupils, reducing external validity for all United States pupils.
In the full sample, covariate adjustment moves the estimated coefficient from to , a large change consistent with substantial measured confounding. In the Catholic-middle-school subsample the estimates remain between and , supporting the claim that restriction has already improved comparability. The cost is visible in the standard errors, which rise from about -- to -- despite similar coefficient magnitudes.
A sufficient causal condition is conditional exchangeability
together with consistency of potential outcomes and positivity in causal inference. For the ordinary-least-squares coefficient itself to equal one common causal effect, also require the correctly specified additive conditional-mean model
Then is the homogeneous treatment effect, and the adjusted coefficient consistently estimates both conditional effects and the average treatment effect in the analyzed population.
Let and let be a consistent estimate. The inverse-probability-weighted estimator of the average treatment effect is
Under the exchangeability, consistency, and positivity conditions in part ii, and consistent estimation of the propensity score, its probability limit is
so it consistently estimates the average treatment effect. It does not require the additive outcome-regression model used to interpret the ordinary-least-squares coefficient.
In the bivariate probit model for endogenous treatment, affects treatment through its threshold equation and affects the outcome through its threshold equation. When , the two disturbances are dependent, so treatment status carries information about the latent outcome disturbance even after conditioning on . Consequently
in general, and the no unmeasured confounding assumption fails.
Fix . For an observation with covariates , put
and let be the standard-normal distribution function. The four conditional cell probabilities are
where the first index is and the second is .
Define as any maximizer of the log likelihood
Then is the requested estimator for the fixed sensitivity value .
The correlation measures dependence between two normalized latent disturbances; it is not a scale-free measure of the strength of one physical confounder. Different latent-variable constructions can induce the same while producing different treatment-outcome confounding, and the same omitted cause can produce different after changing thresholds or disturbance scales. Moreover, re-estimating at each value of changes the entire latent model, so the fitted models do not represent one fixed data-generating mechanism with only its confounder strength varied. At the bivariate normal distribution is singular as well. Varying is a model-based sensitivity analysis, but interpreting the interval as an ordered range of strengths of a single unmeasured confounder is therefore logically unjustified.

Articles by others on the same topic (0)

There are currently no matching articles.