Mean square in ANOVA 2026-10-07
A mean square in ANOVA divides a sum of squares in ANOVA by its positive number of statistical degrees of freedom. If the corresponding residual subspace has scalar error covariance , its expectation is . Its scale depends on whether raw observations or group means were projected.
The additive model gives useful evidence for both predictors. To test for no adjusted sex effect, use against . The statistic has a Student t-distribution with 97 statistical degrees of freedom under the normal linear model and the null hypothesis. Here and . Similarly, against a nonzero common slope gives under the null hypothesis, with . Thus, at a fixed height males have a higher fitted mean weight, and within sex greater height is associated with greater weight.
The interaction model adds . Its female height slope is and its male height slope is . Test against using under the null hypothesis, or equivalently . The -value is 0.73139. The additive model is a reasonable parsimonious choice; the data give no evidence for different height slopes between sexes. This does not establish exact equality of population slopes. The additive model also has the higher adjusted coefficient of determination, 0.405 versus 0.3996, and a slightly smaller residual standard error.
Ignoring sex gives a height slope of 0.8769, appreciably above the adjusted slope 0.6211. This is omitted-variable bias in the descriptive regression: a taller group with a larger intercept makes the pooled height association steeper. Algebraically, with ,
The printed estimates imply a positive height-sex covariance. Pooled and adjusted slopes therefore describe different comparisons; the difference is not evidence that adding sex changed a physical effect of height. For predictions, retain both sex and height, check regression diagnostics, and assess performance on new students. The residual standard error near 5.7 remains substantial and the coefficient of determination near 0.42 does not promise precise individual prediction.
A regression factor is a categorical predictor represented by level indicators and contrasts. Treating year as a regression factor permits arbitrary differences between the three yearly means, reflecting conditions such as weather, rather than imposing a linear time trend.
Let be the yield for seed variety , fertiliser level , and year . A treatment-coded two-factor normal linear model with additive year effects is
with independent errors. The unrestricted coefficients are : six parameters for twelve observations. Equivalently use baseline-zero variety, fertiliser, and interaction arrays. There are no year-treatment interactions in this fit.
Dividing each sum of squares by its statistical degrees of freedom gives the completed analysis of variance:
In particular, the missing interaction statistic is . Its null hypothesis is , against : the change in mean yield from high fertiliser is the same for both varieties. Under the null hypothesis and the independent homoscedastic normal linear model, this nested-model F-test has distribution . The printed -value 0.001137 strongly rejects no interaction.
Do not drop fertiliser just because its main-effect ANOVA row has . In this balanced factorial design that row measures a fertiliser effect averaged over varieties; large opposite effects can cancel. The strongly supported interaction requires retaining its associated main effects by regression model hierarchy. Seed and year also have evidence of effects in their ANOVA rows, so this table alone gives no compelling simplification of the fitted terms.
The interaction contrast is , with standard error 0.7946. The fitted high-minus-low fertiliser effect is for variety 1 but for variety 2. The fitted variety-2-minus-variety-1 contrast is under low fertiliser and under high fertiliser. Thus variety 2 performs best under low fertiliser, while high fertiliser benefits variety 1 and reduces variety 2's fitted yield. The marginal fertiliser effect is only .
The four fitted means in the baseline year are for respectively. Add in 2005 and in 2006 to every fitted mean. The 2005 coefficient has under a test of a zero contrast with 2004; the analogous 2006 contrast has . The latter is lack of evidence for a difference, not proof that the yearly means coincide. Standard errors for the combined contrasts require the coefficient covariance matrix. The experiment is small: residual standard error is 0.6882 on only six statistical degrees of freedom. If the same fields were reused, possible within-field dependence and field allocation would need checking before treating the fitted error independence as established.
There are 23 independent vial responses. The intercept-only fit has 22 residual statistical degrees of freedom and deviance 33.631. Adding two day indicators removes two statistical degrees of freedom and reduces the residual deviance by . Adding the two species indicators then reduces it by . Thus the six replacements, in their printed order, are .
The residual deviance statistical degrees of freedom are , , and . These sequential analysis-of-deviance rows compare nested generalized linear models in the displayed order. In an unbalanced design, the first day row does not test day after controlling for species.
The saturated model estimates each Poisson mean separately as , permitting the boundary value zero when . Subtracting fitted from saturated log-likelihood gives the Poisson deviance
with . A zero observation contributes . If the model includes an intercept, the score equation gives , so the summed linear terms cancel; the unsimplified expression applies generally.
For goodness of fit, the null hypothesis is that the specified independent Poisson mean model is correct; the alternative is an unrestricted collection of means. When a quadratic likelihood approximation is valid, for example with fixed and large expected counts, has approximately a chi-squared distribution with statistical degrees of freedom. Reject for a large upper-tail deviance. Merely increasing the number of sparse observations does not make this saturated-model reference reliable: the alternative dimension also increases. With small fitted means, replication-based checks or a parametric bootstrap are more defensible. This qualification concerns the absolute goodness-of-fit deviance; a fixed-dimensional nested likelihood-ratio comparison can have a better chi-squared approximation.
A Poisson model assumes independent counts between rats conditional on time, a common mean at a given time, and conditional variance equal to that mean. Comparable observation units and consistent counting are needed; unexplained rat-to-rat heterogeneity is excluded by this simple specification. The quadratic Poisson regression is
The printed nested analysis of deviance tests against . Its likelihood-ratio statistic is , approximately under the null hypothesis and regular Poisson asymptotics. The reported gives borderline rejection of the log-linear time effect at 5%, conditional on the Poisson assumptions. This is curvature of the log mean, rather than an ordinary quadratic model for the count itself.
There are only three distinct time values, so the intercept, linear term, and quadratic term span every possible set of three log means. This is Poisson mean saturation at three distinct times: the quadratic fit is saturated for the time-group means, which are , , and . A grouped deviance comparison against unrestricted time-group means has zero statistical degrees of freedom: the quadratic functional form cannot be tested at these three times. The individual-observation deviance 24.515 still describes within-group departure from the Poisson distribution, but small means and zeros make the naive goodness-of-fit calibration doubtful. Poisson variation itself can be checked using the replicates, preferably by a parametric bootstrap; it is not rendered untestable by the saturated mean specification.
Overdispersion means conditional variance exceeding the variance assumed by the Poisson mean model. For example, rats may have different latent susceptibility, so if , the law of total variance gives . A quasi-Poisson model instead uses . The Pearson-residual dispersion estimate is , using .
For the conventional quasi-likelihood comparison with estimated dispersion, use
Its upper-tail probability is about 0.096, so the evidence for curvature no longer reaches 5%. The F-distribution is an approximate calibration, not an exact small-sample result for arbitrary overdispersed counts; quasi-likelihood does not supply a full Poisson likelihood. With known dispersion one would instead compare approximately with . The displayed dispersion code uses the undefined name acf.glm2; its fitted values must be those of aom.glm2 for the stated calculation.
Test for both nonreference days, against at least one nonzero day fixed effect, conditional on chocolate. There are two added regression coefficients and residual statistical degrees of freedom. The nested-model F-test statistic is
Under and the normal linear model assumptions, . Its 5% upper critical value is , so do not reject the absence of day effects; the p-value is approximately . Thus the missing row has two statistical degrees of freedom, mean square , and the F-test value above. These observations do not provide evidence that including day improves the chocolate-adjusted mean model. They do not prove that every possible day effect or chocolate–day interaction term is absent.
The Quasi-Poisson regression retains the same conditional mean but permits
This is a mean–variance function specification through quasi-likelihood; it does not assign a full probability distribution to each count. For independent observations, the quasi-score equation is proportional to , so the mean statistical parameter estimates equal those from Poisson regression. The output estimates through the Pearson dispersion estimator, and inflates the standard errors by approximately .
The large residual deviance relative to 728 residual statistical degrees of freedom also signals substantial overdispersion. Daily weather, traffic and other omitted conditions may produce greater count variation than a homogeneous Poisson distribution allows. The Quasi-Poisson regression accounts for that extra marginal variance. It still requires a correct conditional mean and an appropriate independence assumption; a common dispersion parameter alone does not repair serial correlation.
Allow each subject to have a separate baseline and slope through a correlated random-intercept and random-slope model:
The subject random effects are independent of all measurement errors, and subjects are independent. The within-person covariance between times is , with an additional when the same reading is used twice.
Prefer the model with a random slope. The maximum likelihood estimation refits improve twice the log-likelihood by while adding two covariance statistical parameters. Both the Akaike information criterion ( to ) and Bayesian information criterion ( to ) strongly favour it. The printed nominal likelihood-ratio test is overwhelming. The usual reference chi-squared distribution with two statistical degrees of freedom is not an exact regular calibration, since zero slope variance is a boundary and the intercept–slope correlation is then unidentified; a design-specific parametric bootstrap could calibrate it. This qualification does not undermine the substantial descriptive improvement shown by both information criteria.
The preferred fit estimates population mean baseline strength kg and weekly gain kg/week. The between-person baseline standard deviation is kg, and the between-person slope standard deviation is kg/week. Their estimated correlation is weakly positive, giving random-effect covariance about kg/week; its uncertainty is not supplied. The measurement-error standard deviation is kg. The printed instead describes the correlation between the estimated fixed effects, not between the subject random effects. These statistical parameter estimates come from restricted maximum likelihood; the model comparison uses the separate maximum likelihood estimation fits.
The population mean fitted trajectory is kg. Hence
Using the reported slope standard error, an approximate 95% confidence interval for the population mean gain is kg; a Student t confidence interval with 499 statistical degrees of freedom is almost identical.
Successful individuals can improve far more than the average. Their latent ten-week gains have fitted normal distribution
A clear estimate for an upper-performing group uses a between-person slope quantile: the 95th percentile is kg, and the 97.5th percentile is about kg. The upper 5% of fitted underlying gains begin around 73 kg. These are person-to-person performance quantiles, not confidence intervals for the population mean. Identifying the best observed trainee would require that person's data or fitted subject random effects; the aggregate output cannot identify a literal maximum. If performance means the observed difference of endpoint readings, add the measurement-error variance to the latent gain variance.
The preferred cumulative-response logistic model estimates
Its logit link describes conditional smoking-free probability given baseline predictors and prior successful weeks. Holding the other predictors fixed, males have times the female success odds; the reported p-value is , and the approximate 95% odds ratio confidence interval is . Age has estimated odds ratio per additional year, with p-value , giving little evidence for an age association in this fit.
The combined treatment has conditional success odds ratio versus the reference treatment, with approximate 95% confidence interval and p-value . The fitted conditional odds are about 39% higher for the combined treatment. The trial randomization supports treatment comparisons, but conditioning on accumulated post-treatment outcomes means this coefficient is not directly the marginal total treatment effect.
Each previous successful week multiplies current success odds by , with approximate 95% confidence interval and very small p-value . This is strong fitted persistence. It can reflect true state dependence, unobserved heterogeneity, or an omitted calendar-time trend; this fit alone cannot distinguish them. The baseline intercept implies success probability for a reference-treatment female aged zero with no previous success. That age is outside the study's useful interpretation range, so the intercept chiefly anchors the regression. The Bernoulli distribution dispersion parameter is fixed at one, and the residual binomial deviance is 1122 on 995 statistical degrees of freedom. With individual binary outcomes, comparing that binomial deviance mechanically to a chi-squared distribution is not a reliable general goodness-of-fit test.
Under , the seven free covariate coefficients all equal zero. The likelihood-ratio test statistic is
The larger fit has ten free statistical parameters, compared with three in the constant-rate fit, so the reference chi-squared distribution has seven statistical degrees of freedom. Its 95th percentile is . Therefore reject at 5%; the p-value is approximately , and the covariate model is preferred. The two constrained sex coefficients add no free statistical parameters.
For the dementia-to-death arrow, the hazard ratio for higher versus lower education is , with approximate 95% confidence interval . Thus higher education is associated with a roughly 77.5% lower fitted death transition intensity, conditional on diagnosis age; the interval excludes one. This is an observational association and is not a direct multiplicative statement about death probability.
Each additional year of age at diagnosis multiplies the dementia-to-death transition intensity by , with confidence interval . The estimate suggests an increase, but the interval includes one, so this individual age coefficient is not significant at 5%. Sex has hazard ratio one for this transition by the model's imposed constraint. The output supplies no estimated sex effect or test of that constraint for this arrow. The nonzero sex coefficient for must not be mistaken for a sex effect on .
The constant subspace has one statistical degree of freedom. Day contrasts have , and the within-block ANOVA stratum has . Because the orthogonal block design puts all four independent dose contrasts in that last ANOVA stratum, the residual has statistical degrees of freedom.
StratumSourceDegrees of freedom
MeanGrand mean1
Between daysDays4
Within daysDose4
Within daysResidual41
Uncorrected totalAll observations50
The corrected total has 49 statistical degrees of freedom; dose is tested against the within-day residual. The quantitative dose scale also permits the dose component to be split into linear, quadratic, cubic and quartic orthogonal polynomial contrasts, each with one statistical degree of freedom. This is optional and does not assume the response is linear in dose.
Start with six program sequences: , , , , and . This gives two appearances of each program per row and per column. It also balances transitions: each ordered pair of distinct programs occurs five times across all adjacent sessions.
Apply restricted randomization by randomly assigning the six sequences to the six volunteers and independently permuting the three program labels. Using the supplied numbers, rank the first six from smallest to largest: the sequence indices are . Assign these, in order, to volunteers 1 through 6. Rank the next three numbers attached to symbols : their order is . Assign actual program labels to these ordered symbols, so base , base , base . The final ready-to-use randomized row-column design is:
VolunteerWednesday 1Wednesday 2Wednesday 3Wednesday 4Wednesday 5Wednesday 6
1ABCABC
2CABCAB
3BCABCA
4CBACBA
5ACBACB
6BACBAC
Programs retain their real identities after this label permutation. The researcher should follow the table across chronological afternoons; arbitrary column permutations are deliberately excluded so that transition balance survives. Each program receives twelve sessions and occurs twice in every block. Under the additive block model, the analysis of variance has row, column, program and residual statistical degrees of freedom , respectively, besides the grand mean. Transition balance is useful against differential carryover effects, but is not a proof that they are absent.
A completely randomized design uses eligible animals, randomly selecting to receive and assigning the remainder to . Use comparable follow-up and assess the same response variable in both groups. Its advantage is simplicity and freedom from previous-treatment carryover effects; its disadvantage is that between-animal variation enters the residual and can make a treatment contrast imprecise.
A randomized complete block design with matched pairs first forms pairs using pre-treatment characteristics such as initial disease severity, age or breed. Independently choose which animal in each pair receives , with its partner receiving . The experimental units are animals; pairs are blocks. The average within-pair difference estimates the treatment contrast, and positive within-pair similarity can reduce its variance. Its advantage is control of known heterogeneity; its disadvantage is the need for useful matching, with fewer residual statistical degrees of freedom and little gain if the matching variables are uninformative. Do not construct pairs using outcomes observed after assignment.
A two-period crossover design randomly assigns half the animals to sequence and half to . Each animal receives both treatments in separate periods, with a scientifically justified interval between them and the same outcome assessment after each period. Animal blocks remove persistent between-animal differences, while the two sequences balance treatment against period. Its advantage is potentially high precision from within-animal comparisons; its disadvantage is vulnerability to carryover effects, changing disease state and irreversible effects. It is suitable only when comparing the treatments in both periods remains meaningful and residual effects of the first treatment are adequately controlled. A two-period crossover design does not by itself disentangle arbitrary treatment-specific carryover effects.