Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 2 Solution Created 2026-10-03 Updated 2026-10-06
Let be recall for age level , processing level , and replicate . Write for Younger and for Older. With Older and A as the reference levels in a regression factor, the additive two-factor normal linear model under treatment coding iswhere the errors are independent , with common positive variance. The design matrix is fixed; subjects supply independent observations; the specified conditional mean is correct; and the absence of an interaction term means the age difference is constant across processing levels. These are model assumptions, not consequences of random assignment. The treatment allocation supports independence between processing assignment and background characteristics, but age itself was not randomized.
An older subject in A has and , so the estimated mean is words. Its standard error is the intercept's . There are residual degrees of freedom, giving the 95% confidence interval for this mean:This is a mean-response confidence interval, not a prediction interval for a new individual's recall; the latter also includes the new observation's error variance.
The second two-factor normal linear model adds age-by-processing interaction terms:Its ten free mean parameters are equivalent to one mean for each of the ten cells. To compare the fits, the null hypothesis is . The full residual sum of squares is on degrees of freedom. Adding interaction reduces the additive model's residual sum of squares by , so the reduced value is . The nested-model F-test statistic isIts p-value is . Reject additivity and retain the interaction model. The significant interaction means that a single age effect averaged over all processing methods does not adequately summarize the data.
The regression intercept is the older-A mean. The age coefficient is the younger-minus-older contrast specifically in A. The processing coefficients compare B, C, D and E with A specifically among older subjects. The interaction coefficients are differences between the age contrast in each of those methods and its value in A. Adding the relevant coefficients givesThus A, C and especially D favour younger subjects in their estimated means, whereas B and E have little estimated age difference. In the younger group, D has the largest estimated recall, followed by C and A; B and E are much lower. Among older subjects C is largest, D and A are intermediate, and B and E are lowest. Comparing these orders is a description of estimates, not a claim that every pairwise difference is significant.
Each coefficient's printed standard error estimates its uncertainty; its statistic divides the estimate by that error, and its two-sided p-value uses under the relevant zero-contrast null hypothesis. For example, the older B-versus-A and E-versus-A contrasts have and . The C-versus-A and D-versus-A older contrasts have and . The B interaction has , while the D interaction is borderline at and the other two individual interaction tests are not significant at . These individual tests do not override the joint interaction F-test, and multiple comparisons require care. A simple younger-versus-older contrast outside A combines coefficients and must use their estimated covariance, rather than adding their marginal errors.
The residual standard error estimates . The coefficient of determination is the fraction of centered sample recall variation explained by the ten-cell model. The overall with null law tests all nine non-intercept coefficients jointly against zero, giving extremely strong evidence that the cell means are not all equal. The sequential analysis of variance separates age, process, and their interaction; the balanced design makes the main-effect sums orthogonal, whereas the coefficient table uses the specified reference-level contrasts.
To assess pooling into three processing types, retain age-by-type interaction because the previous test rejects additivity. The reduced model has six cell means. Its null hypothesis says A and C have equal means within each age, and B and E have equal means within each age: four restrictions. Compare it with the ten-cell model usingIn fact the supplied fitted cell means and equal cell sizes determine the increase without raw observations. Pooling two ten-person cells with means adds to the residual sum of squares. Here the four differences are , soTherefore and . At , the data do not reject the three-type reduction, though the result is borderline and is not proof that the pooled means are identical. This factor-level pooling test keeps the previously supported interaction structure; reducing the categories and removing all interaction simultaneously would test a different set of restrictions.