The ordinary least squares objective has gradient . The normal equation therefore gives
The Gram matrix is a positive-definite matrix because has full column rank. More explicitly, for any ,
since . This proves that the displayed solution is the unique global minimum.
The fitted values are , where the hat matrix is
The inverse of the symmetric matrix is symmetric, so . Multiplication gives . Thus is the orthogonal projection onto the column space of the design matrix.
The regression residual vector is with . Consequently , , and the residual sum of squares is
By fitted-residual orthogonality, . Using , the cross-covariance matrix is
Both vectors are jointly normal, so they are also independent by independence of uncorrelated jointly normal variables. The zero covariance calculation itself only needs the common error variance and lack of error correlations.
For the first wind fit, let be wind velocity and electrical output. The simple linear regression is , with independent . It has two fitted mean parameters and residual degrees of freedom. The missing analysis of variance entries are therefore velocity degrees of freedom , velocity mean square , and residual degrees of freedom . The residual mean square is , and , consistent with the printed value after rounding.
Because the model includes a regression intercept, the coefficient of determination is the explained sum of squares divided by the total centered sum of squares:
A large coefficient of determination does not rule out a wrong mean function. In the PDF's first residual-versus-fitted plot, residuals are negative at both ends and positive in the middle. This curved pattern agrees with the visibly flattening output-versus-velocity relationship and motivates a polynomial regression with a quadratic term.
The second wind model is , again with independent common-variance normal errors. Its fitted mean is . The line marked (A) is a two-sided Student t-test of against , conditional on retaining the intercept and linear term. Under the null hypothesis,
Its two-sided p-value is . This is far below , so reject a purely linear mean in favour of the quadratic fit. Equivalently, the extra-term nested-model F-test has and null law .
The third fit is a reciprocal-predictor regression, , with estimated mean . It has two mean parameters rather than the quadratic model's three. Its residual standard error is smaller, versus , and its coefficient of determination is larger, versus . The corresponding residual sums of squares are approximately and . The reciprocal fit is preferable on both these fit measures and parsimony. The two models are not nested, so an ordinary extra-term F-test between them is inappropriate. On the common normal-error likelihood, the difference also favours the third model.
Neither lower residual sum of squares nor a higher coefficient of determination establishes adequate assumptions. The quadratic residual-versus-fitted plot removes the original pronounced curvature; the reciprocal plot also has no comparably obvious mean trend. The low-output residuals appear somewhat more spread out, so check scale-location plots and residuals against velocity for heteroscedasticity. Q-Q plots assess normality; residuals against observation order assess serial dependence; regression leverage and Cook's distance identify influential observations. Further cross-validation, replicate observations at comparable velocities, and prediction errors would help choose between their extrapolation behaviours. Both fitted shapes should be judged principally over the observed positive-velocity range.
The statistician first models the age- and gender-dependent mean by linear regression. Its fitted value is . At fixed gender, a year of age is associated with a decrease of BMI units; the group coded gender has a fitted mean units above the group coded at the same age. This question does not state which gender receives which code. The intercept refers to age zero in the reference group, outside a typical adult-patient range, so its substantive interpretation is limited. Both slope Student t-tests have small p-values, and the overall F-test supports an age/gender-dependent mean. The coefficient of determination is only , leaving considerable individual variation to investigate.
The residual diagnostics check whether mean adjustment is plausible and whether a single normal error law is adequate. The residual-versus-fitted plot has no strong curved mean trend, although its smooth curve is not perfectly flat. The scale-location plot suggests some decline in spread as fitted BMI increases, so common residual variance is a working approximation rather than an established fact. The normal Q-Q plot has systematic departures, notably shorter tails than the normal reference, suggesting the residual law is not exactly normal. The regression leverage plot does not display an obviously extreme leverage point, but the marked observations and Cook's distance should still be checked for influence. Independence between patients cannot be diagnosed from these four plots alone.
Next the statistician extracts the regression residuals and fits an intercept-only normal model as a baseline residual distribution. With an intercept in the original ordinary least squares model, the residuals sum to zero by the normal equation; the near-zero fitted residual mean and its are therefore automatic, not new evidence of good fit. The standard error in this refit uses degrees of freedom, whereas the original uses after estimating three regression coefficients. The reported log-likelihood for the single-normal residual model is .
The histogram suggests a shape worth exploring beyond one normal density, with broad shoulders and some asymmetry, without showing unambiguous separated clusters. The statistician fits two- and three-component finite Gaussian mixtures with a common variance by the expectation-maximization algorithm, starting from separated means and positive weights. Successive likelihoods increase and the displayed final iterations stabilize, as expected from EM likelihood monotonicity. This is evidence of numerical convergence from those starts, not proof of a global maximum.
The two-component fit assigns weights , residual means and common variance . The three-component fit assigns weights , residual means and common variance . The code correctly passes the square roots of those variances to dnorm and overlays the resulting weighted densities. Both curves broadly follow the histogram and are very similar. In particular, a two-component mixture need not have two distinct visible modes.
These fits explore residual mixture clustering: groups differ in BMI relative to the same age/gender-adjusted mean, rather than simply in raw BMI. Membership can be summarized by the fitted mixture responsibilities, retaining uncertainty instead of asserting a certain label for each patient. The common-variance assumption is economical, but should be checked; the common regression slope assumption and constancy of mixture proportions across age and gender are also substantive. A residual mixture alone does not establish genuine biological subpopulations; a flexible continuous distribution, missing predictors, nonlinear mean effects or heteroscedasticity could explain a similar marginal shape.
There is also a two-stage residual mixture fitting limitation. Fitted regression residuals are not exactly independent identically distributed observations: even under a normal regression, their covariance matrix is , and estimating the mean adjustment introduces uncertainty. With subjects and three initial regression coefficients, treating them as independent is a useful approximation for exploring shape, but ordinary mixture likelihoods omit the first-stage uncertainty. A joint mixture regression with shared slopes, or a suitable bootstrap of the whole analysis, would give a firmer basis for parameter uncertainty and cluster selection. Examine component membership versus age/gender and repeat the algorithm from multiple starts before making substantive clustering claims.
Let , , , and . The horizontal axis of the residual-versus-fitted plot is , the fitted values, and the vertical axis is , the regression residual, in millimetres. Under the normal linear model, , where is the regression leverage.
In the quantile-quantile plot, the vertical values are the ordered standardized regression residuals
and the horizontal values are corresponding theoretical quantiles of the standard normal distribution, with plotting positions such as . The exact plotting-position convention has little practical effect here. Under Gaussian errors these points should approximately follow a straight line; they are not independent because fitting induces residual correlations.
The main visible concern is curvature in the conditional mean. The red smooth in the residual-versus-fitted plot descends from positive residuals at low fitted values, becomes negative in the middle, and rises again at high fitted values. This suggests the strictly linear time effects may be inadequate. There is no clear monotone widening of the residual scatter, so strong heteroscedasticity is not apparent. The quantile-quantile plot has modest tail deviations and a few labelled observations, but does not show a dramatic departure from normality. These plots cannot establish independence across days or laboratories; check residuals against day, laboratory, and sampling order as well. A large coefficient of determination does not remove the visible mean-model concern.
Check that the chosen mean structure adequately describes the transformed response. Examine the residual-versus-fitted plot and residuals versus each predictor: systematic curvature suggests omitted nonlinear effects or an interaction term. Plot residual spread against fitted values, for example using a scale-location plot, to assess homoscedasticity. Compare standardized residuals with a normal distribution using a quantile-quantile plot, especially for finite-sample Student's t-distribution and F-test inference.
Also check independence using the sampling design and residuals versus time, collection site, or other groups; spatially related photovoltaic systems may have correlated errors that are invisible in a residual-versus-fitted plot. Investigate large standardized regression residuals, high regression leverage, and influential observations using Cook's distance. Check that the design matrix has full rank and that severe multicollinearity is not making estimates unstable. Reassess whether the response transformation and retained predictors improve these diagnostics and prediction; a coefficient p-value alone cannot establish model adequacy. If observations are independent only conditional on site effects, use an appropriate dependence model rather than treating correlated systems as independent replicates.
Write for stirring rate of observation in furnace . The independent random-intercept and random-slope model is
Errors and random effects are independent. The separate R terms (1 | furnace) and (0 + stir | furnace) impose independent random intercepts and random slopes; (1 + stir | furnace) would instead estimate their covariance as well. The estimates are
The corresponding estimated standard deviations are , , and ; they are not additional parameters.
For a furnace with predictor vector , the marginal formulation, obtained by integrating the Gaussian random effects, is
independently across furnaces. In particular,
In stacked notation , , giving the Gaussian linear mixed model marginal law with .
The residual-versus-fitted plot on page 17 has residuals of both signs over the fitted range, with no convincing smooth curvature or clear fan shape. Its quantile-quantile plot is roughly straight in the central region, with some tail departures and a few large negative residuals. There is no decisive visible violation, but investigate those observations. These plots mainly address conditional observation errors; they do not validate the distribution of furnace random effects or independence within furnaces. With only three furnaces, normality and the variance of the random effects are especially difficult to assess.
A residual-versus-fitted plot for the untransformed fit places on the vertical axis and on the horizontal axis. A fan-shaped cloud, with larger residual spread at larger fitted weight, suggests heteroscedasticity and approximately multiplicative errors. A logarithmic transformation may stabilize this spread when the conditional standard deviation is roughly proportional to the mean.
One can also use the scale-location plot of against fitted values, where is a standardized residual. A systematic rise suggests nonconstant error variance. Neither plot proves that taking logarithms is sufficient: diagnostics of the transformed fit should be checked as well, and dependence between repeated chicken measurements is a separate issue.
Regression diagnostics 2026-10-06
Regression diagnostics check assumptions and influential observations in a fitted regression function. Residual-versus-fitted plots can reveal omitted mean structure or nonconstant variance; a quantile-quantile plot checks a specified error distribution. Regression leverage measures unusual predictor configurations, and Cook's distance combines leverage and residual size to measure coefficient sensitivity. Independence may require checking collection order and the study design in addition to residual plots.
Residual-versus-fitted plot Created 2026-10-05 Updated 2026-10-06
A residual-versus-fitted plot displays residuals vertically against predicted responses horizontally. Curvature suggests an incorrect conditional-mean model; a fan-shaped spread suggests heteroscedasticity, while isolated large residuals may indicate unusual observations. It does not alone measure leverage or influence: regression leverage records unusual predictor combinations, and Cook's distance combines leverage and residual size.
Scale-location plot 2026-10-06
A scale-location plot compares the square root of the absolute standardized regression residual with the fitted values. Systematic changes in level or spread suggest heteroscedasticity. It assesses the error scale, while a residual-versus-fitted plot primarily reveals mean patterns as well as changing spread.