Differentiability in quadratic mean 2026-10-07
A dominated statistical path is differentiable in quadratic mean at if its square-root probability density function has an L2 space derivative , with . The score function is . This formulation controls mass near zeros of as well as the derivative on the support of ; a pointwise log-density derivative is insufficient by itself.
Nuisance tangent space 2026-10-07
The nuisance tangent space is the closed linear span of score functions from statistical paths that vary the nuisance parameter while fixing the target statistical parameter. It is a closed subspace of a Hilbert space inside . Removing its component from a parametric score function leaves the information that nuisance variation cannot imitate.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 1 a Solution Created 2026-10-03 Updated 2026-10-07
A statistical path through is a family , defined on an open interval containing zero, with . Write and . Its differentiability in quadratic mean at zero means that there is a square-integrable function such thatThe function is the score function of the statistical path. Thus it is the derivative of the square-root probability density function, multiplied by two and divided by where . The score function is determined only -almost everywhere. The displayed definition also controls any probability mass entering a region where ; a pointwise derivative of the log probability density function alone would not do that. The required derivative is in the square-root-density norm.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 1 c Solution Created 2026-10-03 Updated 2026-10-07
Fix . Let be the nuisance tangent space: the closed linear span in of score functions of statistical paths that vary only the nuisance parameter . By the preceding argument, is contained in the mean-zero L2 space .
Let denote orthogonal projection onto this closed subspace of a Hilbert space. The efficient score and scalar efficient information areThe efficient score is the component of the parametric score function that cannot be reproduced by changing the nuisance parameter. The efficient information is its squared L2 norm; it can be zero, so positivity must not be assumed in the definition.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 2 a Solution Created 2026-10-03 Updated 2026-10-07
Choose a collection of statistical paths through that are differentiable in quadratic mean. A statistical tangent set at is the set of their score functions. In particular each member belongs to , by the mean-zero score identity under quadratic-mean differentiability.
A statistical tangent set records which first-order directions the chosen statistical paths can realize. Its closed linear span in L2 space is the statistical tangent space. The tangent set consists of attainable scores; the tangent space also includes their linear combinations and limits.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 2 b Solution Created 2026-10-03 Updated 2026-10-07
Let be a bounded function with . A bounded density tilt realizes this direction:These are nonnegative probability density functions, since . The uniform Taylor expansion of the square root givesThe squared L2 norm of this remainder is , since is bounded. Thus the statistical path is differentiable in quadratic mean with score function .
Conversely every score function is centered by the mean-zero score identity under quadratic-mean differentiability. Choosing all these bounded density tilts therefore gives the statistical tangent setThis is a valid choice of statistical tangent set; it does not assert that every possible score function in the unrestricted density model is bounded.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 2 e Solution Created 2026-10-03 Updated 2026-10-07
For the bounded density tilt , differentiate the polynomial in :Since every score function is centered, the centered representer isA bounded baseline makes this a bounded mean-zero function, hence an element of . Thus it is the efficient influence function for the density fourth-power functional relative to the regular bounded density tilts. Its squared L2 norm is .
A bounded baseline alone does not make this functional differentiable along every quadratic-mean differentiable path. The derivative above is the intended regular-path answer. To see the need for the qualification, let and put on , extended by zero outside. For , defineEach is a nonnegative continuous probability density function, because its narrow bump has mass . Moreover,It is therefore a differentiable-in-quadratic-mean path with zero score function. YetThe density fourth-power functional is not even continuous along this statistical path, although every is individually bounded. Consequently no efficient influence function represents derivatives over the unrestricted class of all such paths.
One sufficient additional condition is a common bound for all small . Taylor expansion then bounds the fourth-power remainder by , whileTogether with the quadratic-mean to L1 density derivative, this gives . Under this local bound, or when the chosen statistical paths are the bounded density tilts, the boxed canonical gradient is fully justified. The counterexample is a spike obstruction to density-power differentiability.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 4 c Solution Created 2026-10-03 Updated 2026-10-07
Now . With and the location score , differentiation of the log-likelihood givesTake the usual regularity conditions and so that this score function belongs to L2 space.
For the specified nuisance statistical path, differentiating at zero gives the score function . Normalization requires . There is a second constraint: the statistical path must preserve the model's zero error expected value. ThusThis constraint follows from being a path through the model, rather than from normalization alone. For a centered symmetric logistic distribution, for example, is bounded and has zero expected value, but ; it would not preserve the mean.
For any , independence impliesAll these products are integrable by the Cauchy-Schwarz inequality, since is square-integrable. Thus every such error-density score function is orthogonal to every . Covariate-density nuisance score functions are also orthogonal to these functions, since . This is the mean-preserving error tangent space orthogonality used in the final calculation.
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 36 4 d Solution Created 2026-10-03 Updated 2026-10-07
First use the additional admitted form of the efficient score. Write and let . The difference is the orthogonal projection onto the nuisance tangent space. Part (c) makes every orthogonal to that space. Therefore, for all ,The function in braces belongs to L2 space; choosing it as shows it vanishes -almost everywhere. Hence the required conditional deduction isUnder the regular tail condition at both infinities, integration by parts gives , and the formula becomes . This also confirms the sign.
The admitted product form is an extra restriction; it does not follow for every independent-error regression model. To locate the restriction precisely, suppose the error-density nuisance statistical paths preserve both normalization and mean to first order. Their closed mean-preserving error tangent space is . The full nuisance tangent space is the orthogonal sum of this space and the centered functions of .
For completeness, bounded functions satisfying the two constraints are dense in the error space. Truncate an arbitrary element, subtract its expected value, and then subtract a multiple of a fixed bounded centered function with . Such a exists by truncating , since . The correction coefficients tend to zero by the Cauchy-Schwarz inequality, so the corrected truncations converge in L2 space.
Let . With and , the orthogonal projection of onto the error nuisance space is ; its projection onto the covariate nuisance space is zero. Thus the general efficient score in independent-error regression isThe orthogonality of the two terms gives the displayed efficient information. For a normal distribution of the error, , so this reduces to the admitted formula. It also does so when is constant.
A concrete counterexample to the generality of the admission is , with uniform on and independent standard logistic distribution error. Here and , so the actual efficient score is . It cannot equal because is not constant. This score function is already orthogonal to every nuisance score function: its factor is centered against error-only directions, and its factor is centered against covariate-only directions. The requested formula is valid under its stated additional admission, with the general independent-error formula above explaining its limits.
A statistical functional is pathwise differentiable relative to chosen statistical paths if its derivative along every path depends only on that path's score function and defines a bounded linear functional on their statistical tangent space. The Riesz representation theorem expresses this derivative as an L2 inner product with a unique element of the statistical tangent space, the canonical gradient. A path family and the associated derivative remainder conditions must both be specified; a formal derivative along one convenient family does not establish differentiability along all paths.
Statistical functional 2026-10-07
A statistical functional assigns a target value to each probability measure in a statistical model. Examples include an expected value, a quantile, and an integral of a power of a probability density function. Its behavior along statistical paths determines whether first-order influence-function representers are available.