With the normalization used in the question, kernel ridge regression minimizesThe representer theorem and the normal equations give the fitted-value vector
Let and let be the orthogonal projection of onto this space. The reproducing property makes for every . Writing gives
Conditionally on , the covariance matrix of the fitted vector isSince the eigenvalues of are , the average conditional variance isFor ,Taking and adding the supplied squared-bias bound gives
Because minimizes this conditional upper bound, its expected value is at most the expected bound at any deterministic . Using the assumed eigenvalue comparison,
The false discovery rate iswhere is the total number of rejections and is the number of rejected true nulls. The Benjamini-Hochberg procedure orders the p-values and takeswith when the set is empty, then rejects the smallest p-values.
For , apply the modified procedure to with critical valuesand let be its number of rejections. If the full procedure rejects and makes rejections, then , and converselyUnder Assumption A and the super-uniformity of a true-null p-value,Summing over gives .
Increasing any coordinates of cannot increase the number of modified BH rejections. Henceis an increasing set.
Under Assumption B, put and . The contribution of true null is bounded byNowPositive regression dependence and implyThe resulting sum telescopes to at most one. Thus each true null again contributes at most , proving FDR control under Assumption B.
Under , conditional independence givesMultiplication by the measurable sign and the law of total expectation prove the first identity.
WriteThenConditional orthogonality under and the variance bounds in (ii) giveand similarly for the term containing . The Cauchy-Schwarz inequality bounds the last term bywhich is by .
The leading summands are independent, centered, and have variance because . The central limit theorem and consistency of therefore give, by the Slutsky theorem,This is the generalized covariance measure statistic.
Without the null, condition first on . Sincethe term involving vanishes, givingFor the specified alternative, independence and make the right sideThe constant choice gives zero because , so no first-order power is expected. Takinginstead gives , producing asymptotic power.
Put . Comparing the objective at and givesOn , the preceding duality inequality givesThe triangle inequality then yields
When and ,Writing , the exponential Markov inequality and the supplied chi-square moment-generating-function bound giveThe prescribed equation makes this . A union bound over the groups therefore gives
For disjoint nonempty index sets and , partition the mean and covariance conformably. The conditional multivariate normal distribution isIf the sets overlap, the shared coordinates are fixed and this formula applies to .
Partition into the and blocks. The block-inverse formula identifies the displayed conditional covariance withFor , inversion of the two-by-two precision block givesThe denominator is positive, so the conditional covariance vanishes exactly when . A Gaussian pair is independent exactly when it is uncorrelated, proving
Profiling the Gaussian likelihood over gives . Apart from constants and a positive factor, the negative log-likelihood for the precision matrix isIts derivative is , so the unpenalized minimizer is . The Graphical Lasso solveswith some conventions also penalizing the diagonal.
The Lasso minimizesIts Karush-Kuhn-Tucker conditions areSince the columns of are centered, . Taking the inner product of the KKT equation with and usinggives
The assumed restricted eigenvalue condition and giveCanceling one prediction-norm factor and applying the restricted eigenvalue condition again gives
Articles by others on the same topic
There are currently no matching articles.