Generalized additive mixed model 2026-10-05
A generalized additive mixed model supplements smooth population effects with latent random effects. Given these effects, responses follow a specified response family; integrating them out induces dependence between related observations.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 206 3 e Solution Created 2026-10-03 Updated 2026-10-05
There is a genuine syntax error in the printed command: With the default Gaussian family, this generalized additive mixed model isHere across doctors, is an estimated covariance matrix, and the Gaussian errors have variance and are independent of the random effects. The formula
random=list(...) must be a separate argument, following a comma rather than a plus. The intended call ismodel2 <- gamm(npres ~ s(age, bs="cr") + s(wbc, bs="cr"), random=list(doctor=~age))~age includes a random intercept as well as a random slope; it does not force these two effects to be uncorrelated.The random intercept allows different prescribing baselines, and the random slope allows doctor-specific linear modifications to the population age curve. Conditional on these effects the observations are independent under the model, whereas two patients of the same doctor have additional covariance . This accounts for clustered data and avoids treating doctor variation as independent patient-level noise. The population smooth and doctor-specific predictions are different targets. With only four doctors, estimates of the random-effect covariance can be imprecise; the Gaussian support issue for counts also remains. The new model accounts for doctor-specific baselines and age slopes.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 206 4 a Solution Created 2026-10-03 Updated 2026-10-05
The random-slope linear mixed model iswith independence between the random effects and errors. The estimates areIncubators are treated as exchangeable representatives of possible growth environments, so a random slope permits their growth rates to differ while estimating a population mean growth rate. The formula
0+hours deliberately removes a random intercept: larvae are assigned only after hatching, so it is reasonable to assume no incubator-specific baseline at time zero. Random assignment supports this common-baseline assumption in expectation; it does not prove that observed initial sizes or other baseline differences are exactly identical.For a new incubator, no observations are available to estimate its realized random slope, and its mean effect is zero. Under squared-error loss the best prediction integrates over that effect and the residual error:This is the plug-in conditional expectation given the population parameters, rather than a prediction using a fitted effect from one of the four existing incubators. Ignoring parameter-estimation uncertainty, the prediction variance for a new individual is ; it is not just the uncertainty in the population mean.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 206 4 b Solution Created 2026-10-03 Updated 2026-10-05
Let have columns , and let have one column per incubator, with entry in the column for observation 's incubator and zeros elsewhere. Integrating the Gaussian random effects givesConsequently its marginal log-likelihood isThe likelihood-ratio test of against usesFor each covariance, the optimizing coefficients are the generalized least squares estimator . Both fits should use the same marginal likelihood convention, for example fitting the alternative with
update(fly.model, REML=FALSE) and the null as a Gaussian linear model. The displayed alternative REML criterion alone cannot supply .The variance component is constrained to be nonnegative, and zero lies on the boundary. Thus the interior-parameter hypothesis of Wilks theorem fails; the usual approximation is inappropriate. A limit sometimes applies to a single variance component under additional asymptotic conditions, but four incubators do not give a persuasive large-number-of-groups approximation.
Use a parametric bootstrap under the null: estimate there; simulate independent Gaussian errors with that fitted variance at exactly the observed times and grouping labels; refit both models by maximum likelihood for each simulated sample; and calculate . The upper-tail proportion, conventionally , estimates the p-value. Boundary fits with estimated must remain in the simulation distribution. A restricted-likelihood ratio and a correspondingly calibrated bootstrap are another option when the fixed-effect design is identical, but one must not mix ordinary and restricted likelihoods in a single statistic.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 207 1 e Solution Created 2026-10-03 Updated 2026-10-05
A random-effects meta-analysis allows differences in true treatment effects across countries and eligibility criteria. For trial , let be its estimated log odds ratio and its estimated within-trial variance. A usual approximate statistical model isindependently across trials, with independent of sampling errors. Here is the mean true log odds ratio in the population of comparable trials, and is between-trial heterogeneity, not additional sampling error. Therefore , and for a fitted heterogeneity value,One can estimate by restricted maximum likelihood or another justified method; is a plug-in conditional variance and does not fully account for estimating heterogeneity. With only six studies that uncertainty matters. Different eligibility rules motivate this random effect but do not by themselves establish exchangeability or remove bias; known effect modifiers may warrant stratification or regression.
Use leave-one-study-out influence analysis: fit all six trials, then omit Cohen 1989 and refit both and . Report , the corresponding change in the odds ratio, and changes in the confidence interval and heterogeneity. As a diagnostic with heterogeneity held fixed, writing givesThe fitted Cohen weight would be and . A precise but discordant trial can have large influence; refitting assesses additional influence through heterogeneity. The other five trials' data are absent, so a numerical six-trial influence assessment is not identifiable from the displayed table.
Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 216 3 a Solution Created 2026-10-03 Updated 2026-10-05
Write , , and . We keep the model's printed coefficient literally. Its latent log odds can be writtenThe independence of the is conditional on : they are not marginally independent when the random effects are correlated.
Holding the other predictors and the random-effect realization fixed, increasing predictor by one shifts the latent log odds by , so it multiplies the conditional odds of a late departure by . For a binary snow indicator, this compares the snow and no-snow days with other predictors held fixed. This is a conditional odds ratio, not generally the same as a marginal odds ratio after integrating the random effects.
Randomizing allows unmeasured daily conditions to change the success probability beyond the systematic linear predictor. In this logistic-normal regression with autoregressive random effects, let and . The law of total variance givesso the mixture allows overdispersion beyond a binomial distribution with fixed probability. The correlated part also allows related outcomes on nearby days.
The stationary autoregressive process of order one has and . HenceThe term is unstructured daily variability, whereas is the temporally correlated component under the printed parameterization. Larger gives greater persistence: the lag- correlation of the latent log odds is .
The PDF itself prints both the coefficient and the requested combination . They are inconsistent as a marginal variance. If the intended coefficient were , the variance and off-diagonal covariance would instead be and . The prior's argument list also repeats ; its displayed density indicates the intended second variance parameter is .
There is a further substantive issue with that improper prior. After integrating the latent variables, the observed-data likelihood function is positive and continuous at for fixed finite , positive , and . On a compact positive-volume set of those parameters it therefore has a positive lower bound for small . The prior density is , so integrating over gives . Thus the stated joint posterior distribution is improper, an instance of improper posterior from a log-uniform random-effect scale prior. The requested fixed-parameter conditional laws below are nevertheless proper; their existence does not remedy the joint impropriety.