Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 219 3 iii Solution Created 2026-10-03 Updated 2026-10-05
At the fitted hyperparameters, define , , and , with the matrices formed by pairwise kernel evaluations. Conditional independence of observation noise and future latent values givesThe residual has zero covariance with , so the joint normal law makes it independent of . This derives the conditional multivariate normal distribution and the Gaussian process regression posterior:The covariance is a Schur complement and is positive semidefinite. When nonsingular, its joint density is . Repeated phases can make it singular, in which case the displayed normal law is interpreted on its support.
This predicts the true latent light curve, so no future measurement variance is added. For noisy future observations, add their independent noise covariance to . Fixing the hyperparameters omits their estimation uncertainty, as requested; integrating over their posterior would add it. The periodic model extrapolates by phase, even though all prediction times follow the observed times.
Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 219 3 ii Solution Created 2026-10-03 Updated 2026-10-05
Use the periodic covariance functionIts positive semidefiniteness follows by pulling back a Gaussian kernel to the circle . A zero Gaussian process mean is an ensemble statement, whereas an exactly zero long-term average for every periodic light curve is stronger. To impose the latter literally, use a periodic Gaussian process with zero period average, replacing the kernel byHere is a modified Bessel function, and is the period average of the kernel. Subtracting each process's random period average produces this covariance, so it remains positive semidefinite. We use to enforce the stated mean subtraction; using gives the usual ensemble-zero-mean formulation and the same likelihood derivation.
Let and . The prior latent vector is and the independent Gaussian noise is , so with . The Gaussian-process marginal likelihood isPositive measurement variances make positive definite even if is singular. Maximize over all three hyperparameters, using a broad search in to find competing aliases, followed by joint optimization. Cholesky decomposition evaluates the likelihood without explicitly forming an inverse.
For a well-resolved interior maximum, let be the Observed Fisher information in the coordinates . Then . Equivalently, the profile-likelihood curvature in gives the same local nuisance-adjusted variance; keeping the other hyperparameters fixed would generally underestimate it. A profile interval with is a local approximation. If the period likelihood is multimodal or poorly constrained, fit the complete posterior with proper hyperparameter priors and report its modes and marginal credible intervals rather than a single Hessian error bar. The sampling pattern can create period aliases that a single local search misses.