Solution (source code)

= Solution

For the <normal linear model>, the <log-likelihood> is
$$
\ell(\beta,\sigma^2)=-\frac n2\log(2\pi\sigma^2)-\frac{(Y-X\beta)^T(Y-X\beta)}{2\sigma^2}.
$$
The full column rank of the <design matrix> makes $X^TX$ invertible. Differentiating in $\beta$ gives the normal equations $X^TX\widehat\beta=X^TY$. With the minimized <residual sum of squares> substituted, differentiation in $\sigma^2$ gives
$$
\boxed{\widehat\beta=(X^TX)^{-1}X^TY,\qquad\widehat\sigma^2=\frac{RSS}{n}.}
$$
These are the <maximum-likelihood estimators> almost surely; $RSS>0$ almost surely when $\sigma^2>0$ and $n>p$. The linear transformation of the <multivariate normal distribution> yields
$$
\boxed{\widehat\beta\sim N_p\bigl(\beta,\sigma^2(X^TX)^{-1}\bigr).}
$$
Define the <hat matrix> $H=X(X^TX)^{-1}X^T$. The <fitted values> are $\widehat Y=HY=X\widehat\beta$, the <regression residuals> are $e=Y-\widehat Y=(I-H)Y$, and $RSS=e^Te$. The <orthogonal projection> $I-H$ has rank $n-p$ and annihilates $X\beta$. By <Cochran's theorem>,
$$
\boxed{RSS/\sigma^2\sim\chi^2_{n-p},\qquad\widehat\beta\ \text{is independent of }RSS.}
$$
Consequently $E(\widehat\sigma^2)=(n-p)\sigma^2/n$, with <bias> $-p\sigma^2/n$. The <unbiased estimator> is
$$
\boxed{\widetilde\sigma^2=\frac{RSS}{n-p}.}
$$
Its square root, the reported <residual standard error>, estimates $\sigma$; the assertion of unbiasedness applies to the variance, not generally to its square root.

For the paper-strength analysis, let $h_i$ be the measured percentage and $z_i=h_i-\overline h$. Both <normal linear models> use independent errors of common <variance> $\sigma^2$. Their mean functions are $\beta_0+\beta_1z_i$ for `lm1`, and $\beta_0+\beta_1z_i+\beta_2z_i^2$ for `lm2`, with separately fitted coefficients. The reported <residual standard errors> are \b[$11.82$ and $4.42$], respectively.

For `lm1`, the estimated <conditional expectation> at a new percentage $x$ is
$$
\boxed{\widehat m(x)=34.1842+1.7710(x-\overline h).}
$$
The original data mean $\overline h$ must be used for centering the new percentage; its numerical value is not supplied in the excerpt. Because $\sum_i z_i=0$, the two columns of this <design matrix> are orthogonal, and $\operatorname{Cov}(\widehat\beta_0,\widehat\beta_1)=0$. Hence the coefficient <standard errors> in the output give
$$
\boxed{\widehat{\operatorname{Var}}\{\widehat m(x)\}=2.7108^2+(x-\overline h)^2\,0.6478^2.}
$$
This estimates uncertainty in the mean. Predicting an individual future batch would additionally require the new-error variance, estimated by $11.82^2$.

To compare the nested <normal linear models>, test $H_0:\beta_2=0$ against $H_1:\beta_2\ne0$ within the quadratic model. The <Student t-test> statistic is
$$
T=\frac{-0.63455}{0.06179}\approx-10.27,
\qquad T\sim t_{16}\quad\text{under }H_0.
$$
Its two-sided <p-value> is $1.89\times10^{-8}$. Equivalently $T^2$ is a partial <F-test> statistic with null distribution $F_{1,16}$. \b[Reject the linear restriction and prefer `lm2` at the 5% level.] Its residual spread is much smaller; the increase in the <coefficient of determination> from $0.3054$ to $0.9085$ supports the same conclusion, although the test is the relevant complexity-adjusted comparison.

For <regression diagnostics>, inspect <regression residuals> against fitted means and hardwood percentage for omitted curvature, a scale-location plot for nonconstant <variance>, and a <quantile-quantile plot> against the <normal distribution> for departures from the error assumption. Check unusual observations using <regression leverage> and <Cook's distance>, and examine residuals in collection or batch order if dependence is plausible. Independence and a common <variance> require substantive justification as well as these plots; a small <p-value> for a polynomial term does not itself check the error model.