For a convex function , its subdifferential at is
Its elements are the subgradients of at .
Only the th block can have nonzero subgradient coordinates. If , the Cauchy-Schwarz inequality shows that the unique supporting vector is . At zero, the defining inequality is for every block vector , which is equivalent to . Thus
The differentiable loss has gradient . The subdifferential sum rule and part (b) give
where blockwise
Suppose and minimize . Their midpoint is also a minimizer because is a convex function. The group penalty is convex, while the squared Euclidean norm is strictly convex in the fitted value. If , strict convexity would make the midpoint objective strictly smaller than the minimum. Hence , so the fitted values are unique.
Part (d) makes the residual and therefore unique, so is unique. The Karush-Kuhn-Tucker conditions from part (c) imply that every nonzero block satisfies . Hence for .
Any two minimizers have the same fitted value and vanish outside . Their difference is therefore supported on and satisfies . If has full column rank, then , proving uniqueness of .
Optimality gives . Substitute , expand both squared norms, and cancel . Rearranging gives
This exposes a factor-of-two typo in the paper: its requested display has rather than on the left while leaving the right side unchanged. For the objective printed in the paper, the displayed inequality above is the correct basic inequality. Part (b) explicitly asks us to use the stronger stated version, so the subsequent argument proceeds from that requested premise.
Let . Standard sub-Gaussian random variable and Gaussian-product concentration gives, for ,
Thus makes this probability at most . On the complementary event, the basic inequality and Holder inequality give
The parenthesis is at most by the triangle inequality and at most by . Substituting proves , whose probability tends to one.
Put . Since has at most nonzero coordinates, . On ,
where . Write . On ,
Squaring after division by proves
A symmetric function is a positive-definite kernel when for every finite set of points and real coefficients. A Reproducing-kernel Hilbert space is a Hilbert space of functions whose evaluation maps are continuous. Its reproducing kernel satisfies and .
Let be uniform on , and independently let have density
The supplied Fourier transform identity, after the change of Fourier convention, gives
Also . Therefore
so one may take .
For any and , part (b) and linearity of expectation give
Symmetry is immediate, so is a positive-definite kernel.
On a grid with mesh proportional to , Hoeffding inequality and a union bound give
The density of has exponential tails. A Chernoff bound therefore shows that is bounded by an absolute constant except on an event of probability . On that event, both the empirical kernel and are uniformly Lipschitz, so every pair is approximated by its nearest grid pair with total error at most . Enlarging constants and using yields
A p-value for is super-uniform under that null: for every . For the Benjamini-Hochberg procedure, order , set
with if the set is empty, and reject the hypotheses whose p-values are at most .
Let be the number of rejections and the true-null indices. For , remove and let be the number determined by the corresponding leave-one-out step-up rule. On one has and , while is independent of . Hence super-uniformity gives
Summing over proves that the false discovery rate is at most .
Under the intersection null, every rejection is false, so the false-discovery proportion is and its expectation is the familywise error rate. Because the p-values are independent and exactly uniform, every super-uniform inequality in part (b) is an equality. Here , and therefore
Equivalently, the order-statistic identity in the hint gives the same equality by induction on .
The ridge regression estimator minimizes
Its gradient vanishes exactly when . Since makes this matrix positive definite,
Use a singular value decomposition . Then
Each nonzero singular value contributes , while each zero one contributes zero. Thus
using the Moore-Penrose inverse.
Put . Its ridgeless minimum-norm fit is , so its first coordinates are
The strong law of large numbers applied entrywise gives almost surely as . Continuity of inversion then yields
almost surely, where the middle equality is the push-through identity.
Because is independent, conditional prediction risk equals . With and ,
The noise term has conditional mean zero and covariance . Taking squared norms proves
Apply the stated deterministic equivalent for the squared resolvent with . The bias term tends to . For the variance use
Normalized traces and therefore give
Thus the conclusion printed in the paper also has both signs involving reversed. The hypotheses as printed imply the formula above; they cannot imply the requested one because is positive definite while for a resolvent limit.

Articles by others on the same topic (0)

There are currently no matching articles.