Solution (source code)

= Solution

The missing output is needed to identify the actual regression formula, response column, transformations and coefficient estimates. A useful possible analysis regresses the maintained-school percentage among acceptances on the corresponding percentage among applications, allowing for year; it must be described as a model of college-level composition, not a direct estimate of individual acceptance probabilities.

For instance, if $a_{jt}$ and $p_{jt}$ are these two percentages, a <normal linear model> might specify $\mathbb E(a_{jt})=\alpha+\beta p_{jt}+\delta\mathbf1_{\{t=2000\}}$, with a slope-by-year <interaction term> if needed. The intercept is the expected response at zero predictor, often an extrapolation; the slope is a percentage-point change in acceptance composition per percentage-point change in application composition. A slope of one is not by itself evidence of equal acceptance chances: equal chances within every college would give the entire relation $a_{jt}=p_{jt}$, including zero intercept and no year shift. A fit with positive slope mainly confirms that applicant composition helps predict acceptance composition.

<Regression diagnostics> should include residuals versus fitted values and predictors, a residual <normal Q-Q plot>, checks of <heteroscedasticity>, <regression leverage>, studentized residuals and <Cook's distance>. The aggregate mature-college category merits an influence check and a sensitivity analysis: combining several colleges can conceal different within-college patterns. Repeated observations from the same college across years may have correlated errors, so treating them as independent needs justification. Check nonlinearity, the year interaction and the effects of excluding no observation except with a substantive reason.

Percentages are bounded, and their precision depends on their denominators. A simple <weighted least squares> or binomial analysis requires the underlying numbers accepted, not merely percentages. For a proportion based on $m$ people, the working variance is $\pi(1-\pi)/m$ under independent sampling, rather than constant across colleges. If the predictor percentage is itself estimated, an ordinary regression also ignores its uncertainty. The supplied summaries cannot recover all four cells of each college <contingency table> or the actual fitted numerical output.

The https://www.admin.cam.ac.uk/reporter/2000-01/special/07/13.html[published Cambridge percentage table] nevertheless allows a descriptive comparison of acceptance rates. Write $p$ for the maintained-school proportion among applicants and $a$ for its proportion among acceptances. The <relative risk from group compositions> gives
$$
\boxed{\frac{P(\text{accepted}\mid\text{maintained})}{P(\text{accepted}\mid\text{other})}=\frac{a(1-p)}{p(1-a)}.}
$$
The overall acceptance <probability> cancels by <Bayes' theorem>. This <risk ratio> compares acceptance <probabilities>; it is not the acceptance <odds ratio>, which also needs the overall acceptance <probability>.

For Christ's, the two published application/acceptance pairs give <risk ratios> $37/77\approx0.481$ and $147/187\approx0.786$. If the displayed percentages were rounded to the nearest whole percentage point, the <relative risk bounds from rounded group compositions> give ranges approximately $[0.461,0.501]$ and $[0.755,0.818]$. Thus both comparisons remain below one under that rounding assumption. These are descriptive comparisons of observed groups, without adjustment for applicant qualifications or a <causal effect> interpretation. \b[The published compositions determine approximate relative acceptance rates, but not their <standard errors> or the original regression results.]