With the normalization used in the question, kernel ridge regression minimizes
The representer theorem and the normal equations give the fitted-value vector
Let and let be the orthogonal projection of onto this space. The reproducing property makes for every . Writing gives
Conditionally on , the covariance matrix of the fitted vector is
Since the eigenvalues of are , the average conditional variance is
For ,
Taking and adding the supplied squared-bias bound gives
Because minimizes this conditional upper bound, its expected value is at most the expected bound at any deterministic . Using the assumed eigenvalue comparison,
The false discovery rate is
where is the total number of rejections and is the number of rejected true nulls. The Benjamini-Hochberg procedure orders the p-values and takes
with when the set is empty, then rejects the smallest p-values.
For , apply the modified procedure to with critical values
and let be its number of rejections. If the full procedure rejects and makes rejections, then , and conversely
Under Assumption A and the super-uniformity of a true-null p-value,
Summing over gives .
Increasing any coordinates of cannot increase the number of modified BH rejections. Hence
is an increasing set.
Under Assumption B, put and . The contribution of true null is bounded by
Now
Positive regression dependence and imply
The resulting sum telescopes to at most one. Thus each true null again contributes at most , proving FDR control under Assumption B.
Under , conditional independence gives
Multiplication by the measurable sign and the law of total expectation prove the first identity.
Write
Then
Conditional orthogonality under and the variance bounds in (ii) give
and similarly for the term containing . The Cauchy-Schwarz inequality bounds the last term by
which is by .
The leading summands are independent, centered, and have variance because . The central limit theorem and consistency of therefore give, by the Slutsky theorem,
This is the generalized covariance measure statistic.
Without the null, condition first on . Since
the term involving vanishes, giving
For the specified alternative, independence and make the right side
The constant choice gives zero because , so no first-order power is expected. Taking
instead gives , producing asymptotic power.
The Group Lasso penalty is
By the Cauchy-Schwarz inequality within each group,
Put . Comparing the objective at and gives
On , the preceding duality inequality gives
The triangle inequality then yields
When and ,
Writing , the exponential Markov inequality and the supplied chi-square moment-generating-function bound give
The prescribed equation makes this . A union bound over the groups therefore gives
For disjoint nonempty index sets and , partition the mean and covariance conformably. The conditional multivariate normal distribution is
If the sets overlap, the shared coordinates are fixed and this formula applies to .
Partition into the and blocks. The block-inverse formula identifies the displayed conditional covariance with
For , inversion of the two-by-two precision block gives
The denominator is positive, so the conditional covariance vanishes exactly when . A Gaussian pair is independent exactly when it is uncorrelated, proving
Profiling the Gaussian likelihood over gives . Apart from constants and a positive factor, the negative log-likelihood for the precision matrix is
Its derivative is , so the unpenalized minimizer is . The Graphical Lasso solves
with some conventions also penalizing the diagonal.
The Lasso minimizes
Its Karush-Kuhn-Tucker conditions are
Since the columns of are centered, . Taking the inner product of the KKT equation with and using
gives
Put . On ,
Using and
in the basic inequality yields
In particular lies in the Lasso cone condition.
The assumed restricted eigenvalue condition and give
Canceling one prediction-norm factor and applying the restricted eigenvalue condition again gives
Choose
Every null coordinate satisfies . For ,
Thus on .

Articles by others on the same topic (0)

There are currently no matching articles.