The three Poisson regression models have independent counts and a logarithmic link function: their mean predictors include both
pc and urban in m1, only urban in m2, and only pc in m3. Nested comparisons through the full model are preferable to a direct comparison of the two nonnested one-predictor models.To remove
urban from m1, test against . The likelihood-ratio test statistic is the deviance differencewith approximate null distribution . It is below , so do not reject at 5%. The corresponding Wald p-value is consistent with this decision. To remove pc from m1, test against a nonzero coefficient, giving with approximate null distribution ; reject decisively. The Akaike information criterion also prefers m3, whose value is smaller than and .Among these three models choose
m3, retaining building cover and omitting the urban indicator. This is a comparison within the stated Poisson family; its absolute fit must still be checked, as the next part does.The chosen Poisson regression has fitted meanFor zero building cover, its fitted expected number of floods over the 25-year observation window is . Increasing the covered proportion by multiplies the fitted expected count by . Thus a ten-percentage-point increase, , multiplies the expected count by about , an increase of . A one-percentage-point increase multiplies it by about . The full-unit multiplier compares proportions differing by one, not percentages differing by one.
These are associations in the conditional expectation with the recorded final-year building cover. The regression coefficient is not automatically a causal effect of changing development, especially since the count covers 25 years while cover was measured at the end.
The residual deviance is on residual degrees of freedom. Under a suitable Poisson deviance goodness-of-fit test approximation it is compared with , whose supplied 95th percentile is . Since is much larger, the chosen Poisson model does not provide a satisfactory absolute fit, despite being preferred among the three candidates.
The ratio suggests substantial overdispersion or mean-model misspecification. It is a deviance ratio, not the exact Pearson dispersion estimator, which cannot be computed from the excerpt. Check residual patterns, nonlinear effects, unusual rivers and dependence. A Quasi-Poisson regression could adjust uncertainty under an estimated dispersion, and a negative binomial regression could model extra count variation; a more adequate mean model may also be needed. The nominal Poisson tests in part (a) should be interpreted in light of this failure, rather than treated as a validated final analysis.
A 100-year return level is a flow threshold whose probability of being exceeded in a given year is one percent, under a stationary model for annual extremes. Equivalently, its long-run average waiting time between years with an exceedance is one hundred years. It does not imply that exceedances occur regularly one hundred years apart, or that one is guaranteed in any particular hundred-year interval. This definition uses yearly exceedance events even when the underlying observations are daily maxima.
This is the generalized Pareto distribution for the excess , with scale and shape . Its support requires . The Pickands-Balkema-de Haan theorem says that, for distributions in an appropriate extreme-value domain of attraction, the conditional distribution of excesses above a sufficiently high threshold approaches a generalized Pareto form. This makes it the natural peaks-over-threshold method model; it is an asymptotic justification, not an assertion that every threshold is sufficiently high.
For , the tail decays as a power and has no finite upper endpoint, giving a heavy tail. For , the limiting distribution is exponential, with survival probability . For , there is a finite upper endpoint , giving a bounded tail. Thus the sign of the shape parameter distinguishes heavy, exponential-type and bounded tails.
In the original PDF's probability plot, most points lie close to the diagonal. The quantile plot likewise agrees well over most of the range, but the largest empirical flow, around , is well above its fitted quantile, around . The density plot captures the strongly decreasing right-skewed bulk of the excesses. The return level plot shows a broadly reasonable fit to most observations, with one conspicuously high extreme and increasingly wide uncertainty at long return periods. Consequently the generalized Pareto distribution is a reasonable first approximation, but its most extreme tail is less convincingly represented.
Reading the return-level curve at the 100-year tick givesThese are approximate graphical readings, not parameter-based calculations; rounding to whole units would give about with an interval roughly to . The two blue curves are uncertainty bounds for the fitted return level, not prediction bounds for individual observations. The daily-data independence assumption and the stability of the fit as the threshold changes should be checked, for example by examining clusters of high flows and using declustering of extremes where necessary. A 100-year level also extrapolates beyond the 25-year record.
For the geometric distribution supported on positive integers, set . The geometric series and its derivative giveEquivalently, the sum of the tail probabilities is . For , the expected wait is 50 years, so this is the 50-year return level. The exceedance probability determines the return period, but does not by itself determine a numerical flow rate.
Articles by others on the same topic
There are currently no matching articles.