Finite element 2026-10-05
A finite element specifies a reference region, a finite-dimensional space of local trial functions, and degrees of freedom that uniquely determine each such function. Assembling compatible copies over a finite element mesh gives a global trial space. For a one-dimensional linear element, the region is an interval, the functions are affine functions, and the degrees of freedom are the endpoint values.
The starting point of a conforming finite element method is a weak formulation in a Hilbert space . For example, a homogeneous Dirichlet boundary condition on a bounded Lipschitz domain gives . One seeks
where is a bounded bilinear form, , and is a bounded linear functional. A conforming finite element space is finite dimensional and is typically built from functions that are polynomial on each element of a finite element mesh. With a basis , the Galerkin method imposes the same identity only for , giving
The stiffness matrix is sparse when basis supports overlap only locally. Local element integrals assemble its entries; the mass matrix represents an term. This separates the choice of trial functions from the variational principle determining their coefficients.
The fundamental well-posedness theorem is the Lax-Milgram theorem: if for some , there is a unique solution for each , with . The same theorem on gives a unique discrete solution. The constants must be independent of the mesh for uniform stability of a numerical method. Boundary conditions are built into the trial space when they are essential constraints; flux conditions arise through boundary terms in integration by parts and are natural constraints in the weak formulation.
When is symmetric and a coercive bilinear form, the Ritz method minimizes
over . Its first variation is exactly the Galerkin method, so Ritz minimization and Galerkin testing coincide for symmetric coercive problems. The energy norm turns the approximation into an orthogonal projection. Indeed, subtracting the exact and discrete weak equations gives Galerkin orthogonality, . For every the Pythagorean identity becomes
so
This exact best-approximation property is the central advantage of the Ritz method.
Symmetry is unnecessary for the Galerkin method. For any bounded, coercive form, the Céa lemma gives
To prove it, put . Galerkin orthogonality yields , so coercivity and boundedness give . Divide when and take the infimum. Thus approximation capability plus uniform coercivity gives convergence. In contrast, minimizing for a nonsymmetric form differentiates its symmetric part and generally does not solve the original weak equation.
A concrete symmetric example is the one-dimensional reaction-diffusion problem from Question 5. Continuous piecewise-linear hat functions give the diffusion stiffness matrix with diagonal and neighbouring entries , and the mass matrix with diagonal and neighbouring entries . Therefore and . Its positive-definite matrix is the coordinate form of the continuous energy. The exact smooth solution makes this an explicit test of the method, not just an abstract existence result.
For a nonsymmetric example, take positive , constant real , and
Its form is . Since , it has coercivity constant in the full norm, although it is not symmetric when . The Galerkin method remains well posed. On the uniform hat basis, , where , and . The skew-symmetric matrix cancels from , proving nonsingularity. A Ritz energy would omit precisely this advection contribution. Small may nevertheless make the approximation constant large and produce poorly resolved layers; algebraic solvability is not an accuracy guarantee.
Quantitative convergence uses a finite element interpolation estimate. On a shape-regular mesh, for piecewise-linear conforming elements and in the usual low-dimensional setting,
The Céa lemma therefore gives an energy error. The Aubin–Nitsche duality argument improves the L2 norm error when the dual elliptic problem has regularity: solve and use
Consequently
with constants depending on the solution and regularity bounds. The second estimate is conditional on dual regularity; corners or insufficient data regularity can reduce the rate.
The same Galerkin method handles evolution equations. For the homogeneous heat equation, the semidiscrete finite element heat equation is . Because is a positive-definite matrix and is a positive semidefinite matrix,
The decreasing quadratic form is exactly the L2 norm squared of the finite-element function. This separates spatial stability of a numerical method from the subsequent time integrator, just as in the Schrodinger equation example. Conformity, coercivity and approximation estimates together explain convergence.