Let and be the potential outcomes for the apprenticeship proportion at school if it is or is not converted to a UTC. The policy-relevant estimand is the average causal effect
over the population of schools eligible for conversion. Ideally this would be identified by a Randomized controlled trial of conversion, or by an observational design that reproduces its exchangeability.
The raw difference is generally not a causal effect. Existing UTCs may differ systematically from non-UTCs in location, admissions, pupil composition, prior attainment, resources, and the local apprenticeship market. These are possible confounders of school type and outcome. The comparison also weights the 43 UTCs and roughly 3500 controls very differently from a clearly defined target population. Without conditional exchangeability, positivity, and consistency, the difference mixes the effect of UTC status with selection bias.
Matching in causal inference can improve the comparison by balancing measured pre-treatment covariates and restricting it to non-UTCs resembling UTCs. Under conditional exchangeability given the matching variables, adequate common support, and consistent treatment definitions, it can estimate an effect for the matched population.
It does not remove bias from unmeasured or poorly measured factors such as prior apprenticeship culture, catchment-area opportunities, selection of motivated pupils, or pre-conversion trends. Matching percentages also need not balance their nonlinear effects or interactions. The proposal is better than the raw comparison, but its causal interpretation still rests on untestable assumptions.
Using 50 controls per UTC can reduce sampling variance by averaging over more control outcomes. The gain has sharply diminishing returns because controls matched to the same UTC are not 50 independent treated-control contrasts.
The disadvantage is poorer covariate balance: the 50th-nearest school will usually be much less comparable than the fifth-nearest school, increasing residual confounding and changing the target population. A caliper or variable matching ratio should therefore prevent distant matches even if many controls are available.
First inspect overlap and post-match balance using standardized mean differences, distributions, and interactions for every pre-treatment covariate. Material imbalance or UTCs outside the control support shows that the design is extrapolating and cannot justify conditional exchangeability on those variables.
Second perform a negative control outcome analysis using apprenticeship outcomes from before conversion, or another outcome that UTC status could not yet affect. A nonzero estimated effect indicates remaining selection or differential trends. Neither diagnostic proves absence of hidden confounding, so a quantitative sensitivity analysis for unmeasured confounding is also useful.
If sex predicts apprenticeship uptake, extreme UTC sex ratios create both confounding and weak overlap. Matching only the aggregate percentage of boys may leave no genuinely comparable control, and an additive distance can hide a large imbalance in this influential variable. Sex may also be an effect modifier, so the UTC effect can differ across these unusual compositions. One should enforce close sex-ratio balance and estimate sex-specific effects before standardizing them to the intended policy population.
The average treatment effect is more relevant than the average treatment effect on the treated because the contemplated policy changes treatment status for schools that are presently non-UTC, not merely the 43 schools that selected into existing UTC status. More precisely, the ideal target would be the average effect among the candidate schools the minister might convert; if that population is represented by all eligible schools, it is the ATE. An ATT from existing UTCs answers a narrower historical question and may not transport to the policy targets.

Articles by others on the same topic (0)

There are currently no matching articles.