The expected check loss is finite because . It is also Lipschitz in its location argument, with constant at most one. For any fixed , the continuous distribution function makes . The derivative of with respect to , away from that null event, is . Difference quotients are bounded by one, so the dominated convergence theorem gives
The continuous, strictly increasing distribution function has limits zero and one. Its derivative expression changes sign exactly once, from negative to positive, and hence
Equivalently the difference between the expected losses at and is , positive whenever . This proves population quantiles minimize check loss directly.
For the regression argument, define the population risk and its empirical version . Boundedness of and integrability of the error give integrability of and of all these losses. Conditional on , the conditional distribution of the error has unique -quantile zero. Applying the preceding calculation conditionally shows that its expected loss is uniquely minimized when the fitted displacement equals zero. Thus
with strict inequality whenever that displacement is nonzero on an event of positive probability. The stated identifiability assumption therefore makes the unique population minimizer. This is conditional quantile identification.
The Lipschitz property of the check loss gives
Continuity of and its uniform bound, followed by dominated convergence, show that is continuous on .
For the intended estimator constrained to , here is the entire argmin consistency under uniform convergence in probability argument, rather than an invocation of an M-estimator theorem. For a fixed , let . If this set is empty, there is nothing to prove. Otherwise it is compact, and continuity and uniqueness give a strictly positive gap
Let . If minimizes , then
Consequently
The constrained estimator is consistent. Existence of a constrained minimum follows from continuity of the sample criterion and compactness. The argument applies to any measurable choice of minimizer.
There is an actual domain defect in the printed definition: its minimization is over all of , whereas the assumed uniform convergence in probability is only on . That literal unconstrained consistency claim is false. The preceding conclusion needs minimization over , or another assumption that confines the selected minimizers to a set where the convergence and separation arguments apply. No change of question heading is needed to make this qualification explicit.
Here is a counterexample satisfying even identifiability over the whole line. Take a constant covariate, independent standard normal errors, , , , and
This is bounded and continuous, has range , and vanishes only at zero. The conditional error distribution function is continuous and strictly increasing with median zero. The compact-set uniform law of large numbers also holds: the loss is , Lipschitz in , so a finite grid in reduces its uniform difference to finitely many integrable sample averages plus an arbitrarily small grid error. The weak law of large numbers then gives the asserted convergence.
Let be a sample median, and put . Absolute loss over the range is minimized at . Choose the following global empirical minimizer:
Solving the quadratic in shows , so this really minimizes the criterion over the whole line. For odd , symmetry and continuity give . Moreover in probability: for each fixed positive , the proportion of observations below converges to a number greater than one half, and the analogous proportion below converges to a number less than one half. Therefore
along odd . The selected unconstrained estimator is not consistent. This illustrates why a global argmin may escape a compact convergence set.
Use the following normalization for the Lasso estimator, leaving the intercept unpenalized:
A different convention for the factor in front of squared error rescales the regularization parameter. With this convention, centering of the columns gives . Put and . Then and . The Lasso minimizes . A minimum exists since makes this continuous objective coercive; the bound below holds for every minimizer.
We first prove the needed sharp two-sided Gaussian tail bound. If is standard normal and , substituting in its density integral yields
The last integral without is one after multiplication by . In particular this proves the two-sided bound with no extra factor of two.
Since , each has a standard normal distribution. The union bound, which does not require independent columns or independent scores, gives the Gaussian score event for the Lasso
Here the positive tuning choice presupposes and . When the displayed lower bound is negative it is simply a valid, vacuous lower bound. For the prescribed parameter is zero and the probability lower bound is also zero, so no positive-parameter assertion is supplied by that prescription.
Let and . Comparing the Lasso objective at its minimizer and at and expanding the squares gives the Basic inequality for the Lasso
On , the stochastic term is at most . Because , the triangle inequality gives
It follows that
In particular , which is the Lasso cone condition. The given Compatibility condition for the Lasso is therefore applicable and gives .
Adding to the preceding inequality and using compatibility now yields
The last inequality is just the nonnegativity of . Thus , and dropping one nonnegative coefficient-error term gives
This prediction and coefficient error bound for compatible Lasso holds on , and hence with the required probability. If , the earlier basic inequality directly forces on , so the zero right side is also correct.
With the same squared-error normalization as in Question 2, the ridge estimator solves
The unpenalized intercept in ridge regression is because the columns of are centered. Differentiating the remaining quadratic objective gives . Since , the left matrix is positive definite even when is rank-deficient. Thus the closed-form ridge regression estimator is
The second equality again uses centering. If the penalty convention is with an unnormalized squared-error term , replace by throughout; the limiting assertion is unchanged.
For the full-column-rank case, take a singular value decomposition with positive singular values . The ordinary least squares estimator, which is the slope maximum-likelihood estimator in this normal linear model, and the ridge regression estimator are respectively
Consequently principal-component shrinkage by ridge regression multiplies the least-squares coordinate in direction by . Ridge shrinks most strongly in directions with the smallest singular values. These are combinations of predictors that the data distinguish least well.
Near multicollinearity creates small singular values, and the factor in ordinary least squares greatly amplifies noise in those directions. Its coefficient variance there is , while the ridge variance is
which is smaller. The price is bias: the expectation of the ridge coordinate is times the corresponding true coefficient. This bias-variance tradeoff can substantially reduce estimation and prediction error when the poorly identified directions do not contain an excessively large signal. It is not a guarantee of improvement for every coefficient vector. Adding to every eigenvalue also improves the condition number of the normal matrix and makes numerical inversion more stable.
To handle every rank and every relation between and , put and use its spectral theorem for real symmetric matrices. Choose an orthonormal eigenbasis with eigenvalues , and put . Then
If , then , so and . Thus these terms are exactly zero for every positive ; there is no divergent component in the null space. Each remaining term converges to . By the definition of the Moore-Penrose pseudoinverse,
and therefore
This proves the vanishing-penalty ridge limit without an invertibility assumption. The limit is the minimum-norm least-squares solution: every other least-squares solution differs by a vector in , orthogonal to the displayed solution, and therefore has at least as large a Euclidean norm.
Let be the indices of the true null hypotheses, let , and let count false rejections. The familywise error rate is
The Bonferroni correction rejects a null exactly when its p-value is at most . The union bound gives
This controls the familywise error rate for every configuration of true and false nulls, with arbitrary dependence between their p-values. Marginal super-uniformity, , would also suffice.
For the Holm step-down procedure, inspect the ordered p-values in ascending order and stop at the first failed comparison. If the first comparison fails, reject none; this is the convention when the defining set is empty. Ties can be ordered by any fixed rule.
If there can be no false rejection. Otherwise let be the rank of the first true null. Since at most false nulls precede it, and thus . If any true null is rejected, the step-down rule must have rejected this first true null, which requires
Using the union bound on the true-null p-values proves
This first true null argument for Holm control also requires no independence. Holm controls the familywise error rate at level under arbitrary dependence.
Now let count all rejections. The false discovery rate is the expected proportion of rejections that are false, with zero assigned when nothing is rejected:
It differs from the familywise error rate: several false rejections can still represent a small proportion of a large collection of discoveries.
The BH procedure is a step-up rule. Set
If , reject nothing; otherwise reject all hypotheses with . Exactly hypotheses are rejected: if more than p-values were below that threshold, the next ordered value would also satisfy its own larger threshold, contradicting maximality. Thus . Unlike the step-down rule, a failed early comparison does not make this procedure stop.
For the proof, assume that each true-null p-value is uniform on and independent of the entire vector of the other p-values. Joint independence of all p-values is a sufficient condition; the false-null marginal distributions can be arbitrary. Mere uniformity of the true-null marginals without a dependence condition is insufficient for this argument or for general unmodified BH control.
Fix a true null . Replace its p-value by zero and let be the number of rejections made by the Benjamini-Hochberg procedure on the modified vector. This variable depends only on the other p-values, and . The Benjamini-Hochberg leave-one-out identity is
To prove it, suppose is rejected with . Decreasing its p-value to zero leaves every ordered value above rank unchanged, since it was already among the first values. Those higher ranks still fail their thresholds, while rank still succeeds; hence . Conversely, suppose and . Restoring still leaves at least values at most , so . Increasing one p-value cannot increase the maximal successful rank, so . This gives equality and rejection of .
Independence and uniformity now give
Summing over the true nulls proves the exact false discovery rate under independent null p-values:
If the independent true-null p-values are only super-uniform, the same calculation gives the inequality instead of equality. The distinction between the two types of error control and the dependence conditions is essential: Bonferroni and Holm have the preceding guarantees without independence, while the BH proof here explicitly uses it.

Articles by others on the same topic (0)

There are currently no matching articles.