With vanishing moments, a wavelet annihilates every polynomial of degree below . A Taylor polynomial then explains small wavelet coefficients on smooth parts of a signal, and efficient best N-term approximation. Compact support localizes coefficients near a feature, limits the number of boundary interactions, and permits a finite filter implementation. Increasing the number of vanishing moments while keeping an orthonormal basis generally requires a larger support of a function: a finite orthonormal filter with vanishing moments needs at least taps, and the minimal-support Daubechies wavelet has support length in the standard normalization.
The Haar wavelet, , has one vanishing moment, unit support length and discontinuities. It is inexpensive and particularly suitable for piecewise constant data with sharp jumps. A Daubechies wavelet of a larger order has more vanishing moments and a longer finite filter; sufficiently large orders also provide greater regularity. It is useful when smooth trends should yield small coefficients, although the wider support of a function can spread a jump across more coefficients. Choose Haar for compact jump localization; choose a higher-order Daubechies wavelet for smooth polynomial structure. More vanishing moments alone does not make every low-order wavelet highly differentiable.
Write . A linear N-term approximation fixes the indices independently of , normally the first in a prescribed ordering:
A best N-term approximation chooses the indices using : retain coefficients of largest absolute value, resolving ties arbitrarily, and set . The Parseval identity shows why this choice is optimal:
For any fixed index set, the orthogonal projection coefficients minimize the error; optimizing the set then means discarding the smallest squared coefficients. Thus “linear” requires a specified ordering, while “nonlinear” refers to the data-dependent selection.
Use the usual localized, compact support construction of an interval-adapted wavelet basis, including boundary wavelets with the stated vanishing moments, and order the linear N-term approximation by increasing resolution. Also interpret a piecewise polynomial function as having finitely many pieces. These conventions matter: regularity and vanishing moments alone, or an arbitrary enumeration, do not establish the asserted rates.
At scale , a wavelet whose support lies in one polynomial piece has zero coefficient because . Only a bounded number of wavelets per scale can meet a partition point. Their norms are bounded by , and is bounded. Thus and
Retain the fixed number of coarse scaling function coefficients and every nonzero coefficient through level . This uses at most terms, leaving squared error at most . The optimal best N-term approximation is no worse; choose proportional to to obtain for some . In contrast, retaining all wavelets through level costs terms. Choosing the last complete level before gives
These are squared errors; the corresponding errors are and .
For a function smooth on finitely many closed pieces with finite one-sided derivatives, integration by parts gives its Fourier series coefficients, for ,
where includes the jump of the periodic extension at the endpoint when necessary. In particular . Keeping frequencies , with , gives squared error .
A best N-term approximation cannot improve this worst-case rate. For example, has on odd nonzero and zero on even nonzero . Even the optimal selection therefore leaves a squared tail comparable to . Thus for the class with jumps,
or in the norm. Individual functions with no jumps, or additional cancellation, can converge faster. “Smooth except at finitely many points” must include controlled one-sided smoothness; smoothness merely on open pieces permits pathological behavior near their endpoints.
Use localized tensor-product wavelets and finitely many bounded polynomial pieces, with the boundary wavelets adapted as in part (b). A rectifiable smooth curve of length meets dyadic squares of side : subdividing an arclength parametrization into pieces of length at most covers it by that many balls, each meeting only a bounded number of squares. Enlarging squares by the fixed support diameter preserves the count.
Each normalized two-dimensional wavelet has norm . A coefficient meeting the curve is therefore , and the total squared energy of these coefficients at level is . If , all other coefficients vanish by the vanishing moments. Keeping the curve coefficients through level costs and leaves squared error .
The printed part (e) does not repeat . The bound still holds for any fixed polynomial degree when : on a smooth piece, a Taylor polynomial in the variable carrying a wavelet gives coefficient size . There are such coefficients, so their squared energy is . Retain all coefficients through , and curve coefficients through . The cost is and the omitted squared energy is
The best N-term approximation is at least as good as this selection, proving
This argument covers the unqualified finite-degree clause without adding an unnecessary restriction .
A piecewise polynomial function agrees with a polynomial on each member of a partition. Approximation claims usually require finitely many pieces with controlled interfaces. In one dimension, finitely many partition points and localized wavelets with more vanishing moments than the polynomial degree leave only a bounded number of nonzero coefficients per scale, yielding exponential squared best N-term approximation error.
A finite-length smooth curve meets supports of localized two-dimensional wavelets at scale . For a bounded function, their coefficients are , so their total squared energy at that scale is . Keeping these coefficients through level uses terms and leaves squared error . Polynomial cancellation, or adequate approximation of the remaining smooth regions, therefore gives best N-term approximation squared error .