The expected check loss is finite because . It is also Lipschitz in its location argument, with constant at most one. For any fixed , the continuous distribution function makes . The derivative of with respect to , away from that null event, is . Difference quotients are bounded by one, so the dominated convergence theorem gives
The continuous, strictly increasing distribution function has limits zero and one. Its derivative expression changes sign exactly once, from negative to positive, and hence
Equivalently the difference between the expected losses at and is , positive whenever . This proves population quantiles minimize check loss directly.
For the regression argument, define the population risk and its empirical version . Boundedness of and integrability of the error give integrability of and of all these losses. Conditional on , the conditional distribution of the error has unique -quantile zero. Applying the preceding calculation conditionally shows that its expected loss is uniquely minimized when the fitted displacement equals zero. Thus
with strict inequality whenever that displacement is nonzero on an event of positive probability. The stated identifiability assumption therefore makes the unique population minimizer. This is conditional quantile identification.
The Lipschitz property of the check loss gives
Continuity of and its uniform bound, followed by dominated convergence, show that is continuous on .
For the intended estimator constrained to , here is the entire argmin consistency under uniform convergence in probability argument, rather than an invocation of an M-estimator theorem. For a fixed , let . If this set is empty, there is nothing to prove. Otherwise it is compact, and continuity and uniqueness give a strictly positive gap
Let . If minimizes , then
Consequently
The constrained estimator is consistent. Existence of a constrained minimum follows from continuity of the sample criterion and compactness. The argument applies to any measurable choice of minimizer.
There is an actual domain defect in the printed definition: its minimization is over all of , whereas the assumed uniform convergence in probability is only on . That literal unconstrained consistency claim is false. The preceding conclusion needs minimization over , or another assumption that confines the selected minimizers to a set where the convergence and separation arguments apply. No change of question heading is needed to make this qualification explicit.
Here is a counterexample satisfying even identifiability over the whole line. Take a constant covariate, independent standard normal errors, , , , and
This is bounded and continuous, has range , and vanishes only at zero. The conditional error distribution function is continuous and strictly increasing with median zero. The compact-set uniform law of large numbers also holds: the loss is , Lipschitz in , so a finite grid in reduces its uniform difference to finitely many integrable sample averages plus an arbitrarily small grid error. The weak law of large numbers then gives the asserted convergence.
Let be a sample median, and put . Absolute loss over the range is minimized at . Choose the following global empirical minimizer:
Solving the quadratic in shows , so this really minimizes the criterion over the whole line. For odd , symmetry and continuity give . Moreover in probability: for each fixed positive , the proportion of observations below converges to a number greater than one half, and the analogous proportion below converges to a number less than one half. Therefore
along odd . The selected unconstrained estimator is not consistent. This illustrates why a global argmin may escape a compact convergence set.