For the normal linear model, the log-likelihood is
The full column rank of the design matrix makes invertible. Differentiating in gives the normal equations . With the minimized residual sum of squares substituted, differentiation in gives
These are the maximum-likelihood estimators almost surely; almost surely when and . The linear transformation of the multivariate normal distribution yields
Define the hat matrix . The fitted values are , the regression residuals are , and . The orthogonal projection has rank and annihilates . By Cochran's theorem,
Consequently , with bias . The unbiased estimator is
Its square root, the reported residual standard error, estimates ; the assertion of unbiasedness applies to the variance, not generally to its square root.
For the paper-strength analysis, let be the measured percentage and . Both normal linear models use independent errors of common variance . Their mean functions are for lm1, and for lm2, with separately fitted coefficients. The reported residual standard errors are and , respectively.
For lm1, the estimated conditional expectation at a new percentage is
The original data mean must be used for centering the new percentage; its numerical value is not supplied in the excerpt. Because , the two columns of this design matrix are orthogonal, and . Hence the coefficient standard errors in the output give
This estimates uncertainty in the mean. Predicting an individual future batch would additionally require the new-error variance, estimated by .
To compare the nested normal linear models, test against within the quadratic model. The Student t-test statistic is
Its two-sided p-value is . Equivalently is a partial F-test statistic with null distribution . Reject the linear restriction and prefer lm2 at the 5% level. Its residual spread is much smaller; the increase in the coefficient of determination from to supports the same conclusion, although the test is the relevant complexity-adjusted comparison.
For regression diagnostics, inspect regression residuals against fitted means and hardwood percentage for omitted curvature, a scale-location plot for nonconstant variance, and a quantile-quantile plot against the normal distribution for departures from the error assumption. Check unusual observations using regression leverage and Cook's distance, and examine residuals in collection or batch order if dependence is plausible. Independence and a common variance require substantive justification as well as these plots; a small p-value for a polynomial term does not itself check the error model.
Let be speed and the tool type. Each fit is a normal linear model with independent errors. The three mean specifications are
The constraints are treatment coding with type 1 as reference. The parameter counts are , respectively. Model 2 gives parallel lines; model 3 allows interaction terms and distinct slopes.
In the analysis of variance interaction row, adding three independent slope differences costs three degrees of freedom. Its extra sum of squares is the reduction in residual sum of squares, . Its mean square is , and its F-test statistic is
Thus the four missing entries are , , , and , to the precision allowed by the rounded output. This tests against the alternative that at least one slope difference is nonzero. Under and the normal linear model assumptions, the statistic has distribution . Its p-value gives no reason to reject at 5%.
The simpler common-line model is inadequate compared with the parallel-line model. Testing against at least one nonzero type effect, while retaining speed, gives the partial F-test
Its p-value is far below , so reject the common-line restriction. This reduction differs from the sequential type sum of squares in the displayed table, because that table adds type before speed. Recommend model 2: different intercepts and a common decreasing slope.
The selected coefficient estimates give
The intercept estimates the expected lifetime of type 1 at speed zero; if zero speed is outside the data range, it is only an extrapolated intercept. Its standard error is , and its Student t-test ratio is , with two-sided p-value . At any fixed speed, type 2 differs from type 1 by hours, with standard error , and p-value , giving little evidence of a difference. Types 3 and 4 exceed type 1 by and hours, with standard errors and ; their test ratios and and p-values and support positive differences. Each of these individual coefficient tests has null value zero, alternative nonzero, and null distribution .
The speed coefficient means a reduction of hours per extra revolution per minute, or hours per additional 100 rpm, for every type. Its standard error is ; the ratio has null distribution when the common slope is zero, and two-sided p-value .
The residual standard error estimates the common noise standard deviation in hours using 15 residual degrees of freedom, since . The coefficient of determination means that about of the corrected lifetime variation is explained by this fit. The overall F-test statistic tests the simultaneous restriction against at least one nonzero coefficient; its null distribution is and its p-value rejects an intercept-only mean. These conclusions remain conditional on suitable regression diagnostics for independent, homoscedastic, approximately normal errors.
The required sketch plots the four fitted lines. Its speed interval is illustrative because the observed speeds are not supplied; the ordering and vertical gaps are determined by the fitted coefficients.
Figure 1.
Parallel fitted tool-lifetime lines for the four tool types
.