The ordinary least squares objective has gradient . The normal equation therefore gives
The Gram matrix is a positive-definite matrix because has full column rank. More explicitly, for any ,
since . This proves that the displayed solution is the unique global minimum.
The fitted values are , where the hat matrix is
The inverse of the symmetric matrix is symmetric, so . Multiplication gives . Thus is the orthogonal projection onto the column space of the design matrix.
The regression residual vector is with . Consequently , , and the residual sum of squares is
By fitted-residual orthogonality, . Using , the cross-covariance matrix is
Both vectors are jointly normal, so they are also independent by independence of uncorrelated jointly normal variables. The zero covariance calculation itself only needs the common error variance and lack of error correlations.
For the first wind fit, let be wind velocity and electrical output. The simple linear regression is , with independent . It has two fitted mean parameters and residual degrees of freedom. The missing analysis of variance entries are therefore velocity degrees of freedom , velocity mean square , and residual degrees of freedom . The residual mean square is , and , consistent with the printed value after rounding.
Because the model includes a regression intercept, the coefficient of determination is the explained sum of squares divided by the total centered sum of squares:
A large coefficient of determination does not rule out a wrong mean function. In the PDF's first residual-versus-fitted plot, residuals are negative at both ends and positive in the middle. This curved pattern agrees with the visibly flattening output-versus-velocity relationship and motivates a polynomial regression with a quadratic term.
The second wind model is , again with independent common-variance normal errors. Its fitted mean is . The line marked (A) is a two-sided Student t-test of against , conditional on retaining the intercept and linear term. Under the null hypothesis,
Its two-sided p-value is . This is far below , so reject a purely linear mean in favour of the quadratic fit. Equivalently, the extra-term nested-model F-test has and null law .
The third fit is a reciprocal-predictor regression, , with estimated mean . It has two mean parameters rather than the quadratic model's three. Its residual standard error is smaller, versus , and its coefficient of determination is larger, versus . The corresponding residual sums of squares are approximately and . The reciprocal fit is preferable on both these fit measures and parsimony. The two models are not nested, so an ordinary extra-term F-test between them is inappropriate. On the common normal-error likelihood, the difference also favours the third model.
Neither lower residual sum of squares nor a higher coefficient of determination establishes adequate assumptions. The quadratic residual-versus-fitted plot removes the original pronounced curvature; the reciprocal plot also has no comparably obvious mean trend. The low-output residuals appear somewhat more spread out, so check scale-location plots and residuals against velocity for heteroscedasticity. Q-Q plots assess normality; residuals against observation order assess serial dependence; regression leverage and Cook's distance identify influential observations. Further cross-validation, replicate observations at comparable velocities, and prediction errors would help choose between their extrapolation behaviours. Both fitted shapes should be judged principally over the observed positive-velocity range.
Let be recall for age level , processing level , and replicate . Write for Younger and for Older. With Older and A as the reference levels in a regression factor, the additive two-factor normal linear model under treatment coding is
where the errors are independent , with common positive variance. The design matrix is fixed; subjects supply independent observations; the specified conditional mean is correct; and the absence of an interaction term means the age difference is constant across processing levels. These are model assumptions, not consequences of random assignment. The treatment allocation supports independence between processing assignment and background characteristics, but age itself was not randomized.
An older subject in A has and , so the estimated mean is words. Its standard error is the intercept's . There are residual degrees of freedom, giving the 95% confidence interval for this mean:
This is a mean-response confidence interval, not a prediction interval for a new individual's recall; the latter also includes the new observation's error variance.
The second two-factor normal linear model adds age-by-processing interaction terms:
Its ten free mean parameters are equivalent to one mean for each of the ten cells. To compare the fits, the null hypothesis is . The full residual sum of squares is on degrees of freedom. Adding interaction reduces the additive model's residual sum of squares by , so the reduced value is . The nested-model F-test statistic is
Its p-value is . Reject additivity and retain the interaction model. The significant interaction means that a single age effect averaged over all processing methods does not adequately summarize the data.
The regression intercept is the older-A mean. The age coefficient is the younger-minus-older contrast specifically in A. The processing coefficients compare B, C, D and E with A specifically among older subjects. The interaction coefficients are differences between the age contrast in each of those methods and its value in A. Adding the relevant coefficients gives
Thus A, C and especially D favour younger subjects in their estimated means, whereas B and E have little estimated age difference. In the younger group, D has the largest estimated recall, followed by C and A; B and E are much lower. Among older subjects C is largest, D and A are intermediate, and B and E are lowest. Comparing these orders is a description of estimates, not a claim that every pairwise difference is significant.
Each coefficient's printed standard error estimates its uncertainty; its statistic divides the estimate by that error, and its two-sided p-value uses under the relevant zero-contrast null hypothesis. For example, the older B-versus-A and E-versus-A contrasts have and . The C-versus-A and D-versus-A older contrasts have and . The B interaction has , while the D interaction is borderline at and the other two individual interaction tests are not significant at . These individual tests do not override the joint interaction F-test, and multiple comparisons require care. A simple younger-versus-older contrast outside A combines coefficients and must use their estimated covariance, rather than adding their marginal errors.
The residual standard error estimates . The coefficient of determination is the fraction of centered sample recall variation explained by the ten-cell model. The overall with null law tests all nine non-intercept coefficients jointly against zero, giving extremely strong evidence that the cell means are not all equal. The sequential analysis of variance separates age, process, and their interaction; the balanced design makes the main-effect sums orthogonal, whereas the coefficient table uses the specified reference-level contrasts.
To assess pooling into three processing types, retain age-by-type interaction because the previous test rejects additivity. The reduced model has six cell means. Its null hypothesis says A and C have equal means within each age, and B and E have equal means within each age: four restrictions. Compare it with the ten-cell model using
In fact the supplied fitted cell means and equal cell sizes determine the increase without raw observations. Pooling two ten-person cells with means adds to the residual sum of squares. Here the four differences are , so
Therefore and . At , the data do not reject the three-type reduction, though the result is borderline and is not proof that the pooled means are identical. This factor-level pooling test keeps the previously supported interaction structure; reducing the categories and removing all interaction simultaneously would test a different set of restrictions.
The generalized cross-validation curve reaches its minimum at among the supplied grid values. This balances improved conditioning against shrinkage bias using an estimate of prediction error, rather than choosing the penalty with the smallest training residual sum of squares. Nearby values have similar errors, so the plot does not establish a highly precise optimal penalty.
Reading the PDF table at gives the fitted regression intercept and slopes:
These table entries are necessary because the TeX stores the table only inside a figure. The ridge regression slopes are shrunk relative to the zero-penalty fit; their small nonzero values do not represent variable exclusion.
There is a minor source inconsistency: the prose says the predictors are centered, but the table's intercept varies with . With exactly centered predictor columns and an unpenalized intercept it would remain . The numbers above faithfully report the printed table, while part (a) gives the centered formula and the general uncentered conversion. The table therefore reflects a different or incompletely described preprocessing convention.
At fraction , the Lasso regularization path lies between the vertical lines marked steps 4 and 5. In that interval the first five predictors have nonzero slopes: are positive and is negative. The sixth predictor remains at zero until the later part of the path, beyond this fraction. Therefore the chosen Lasso model includes
The regression intercept is retained. In particular the negative trace has already left zero at fraction ; a small slope is not the same as a zero slope. The selected fraction came from ten-fold K-fold cross-validation, and is not a numerical value of the ridge penalty in the preceding parts.
Write . Minimizing the residual sum of squares first over the coefficients of reduces the problem to minimizing . Full column rank ensures . Thus the partial regression formula and its variance are
This uses and the isotropic error covariance matrix .
Since the span of is contained in the column space of , orthogonal projection onto the latter removes at least as much squared norm. Consequently
and inversion proves the requested lower bound. Full rank makes the denominator positive.
Adding to column leaves the column space unchanged. Every old fitted vector is reproduced by keeping all slopes unchanged and replacing the regression intercept by . Uniqueness of the ordinary least squares estimators then proves that slope is unchanged. In particular all columns can be centered by such changes. Applying the previous projection argument to the centered design proves the bound with and ; the normalized inner product is now the usual sample correlation.
Each reported coefficient test is the two-sided Student's t-test of against , conditional on the other predictors. Under the null, has a Student's t-distribution with degrees of freedom. At five percent the intercept is significant, but neither individual test slope is. The final F-test compares the full model against an intercept-only model: . It has degrees of freedom and rejects at five percent, since its -value is . Thus the predictors jointly explain variation, without strong evidence separating their individual conditional effects.
The two test predictors have sample correlation , although each has only moderate correlation with the final exam response. The variance inflation factor is . Their nearly shared direction produces substantial multicollinearity, inflating the individual slope standard errors. The joint F-test can detect this common predictive direction even though either coefficient, given the other, has a large -value.
In ridge regression, penalize slopes but ordinarily leave the regression intercept free. Centering gives and . An exactly centered predictor matrix has for every penalty. A table with a penalty-dependent intercept therefore cannot arise from exactly centered columns under this convention.