Put and . The density can be writtenwhereThus this is an exponential dispersion family. Its derivative identities giveThe variance function is . In the usual Gamma-GLM convention the canonical inverse link isit differs by a minus sign from the natural parameter .
Let be row of the design matrix, , and . The full log-likelihood isDifferentiation gives the score functionandThe Hessian does not depend on , so the Fisher information matrix is
The score equation is independent of , because is only a common factor. Fixing or estimating it therefore gives the same maximum-likelihood estimator .
Under the usual full-rank and regularity conditions,with evaluated consistently at the fitted means. Consequently every standard error from "mod2" is times the corresponding standard error computed with dispersion one in "mod1". In particular,
The null hypothesis isagainst the alternative that at least one coefficient is nonzero. The code uses the scaled reduction in exponential-family deviance,and compares it with .
That chi-squared calibration is appropriate when the dispersion is known, and is asymptotically valid after consistent dispersion estimation. With unknown Gamma dispersion, the standard finite-sample GLM comparison instead usesagainst an distribution. Its p-value is approximately , so the correctly calibrated test still rejects at the 5% level and gives evidence that component type or position contributes to mean failure time.
First recode the labels as ; with labels and , the displayed constraints for class zero can never hold. Add an unpenalized intercept so that the separating hyperplane need not pass through the origin, and solveCentering and scaling predictor columns is also advisable because the Euclidean penalty depends on their units.
The support vector machine classifier isUnder the constraints, the two classes lie beyond the parallel support-vector-machine margin boundaries . Their distance is , so minimizing the norm maximizes the geometric margin. The observations touching the margin are the support vectors.
The hard-margin constraints are feasible exactly when the classes are separated by a separating hyperplane. Failure after the corrections in part a therefore means the classes overlap.
Introduce slack variables of a support vector machine and solve the soft-margin problemsubject toLarge strongly penalizes violations and approaches the hard-margin solution when separation is possible. Small tolerates more violations in exchange for a wider, more strongly regularized margin.
The kernel trick replaces inner products by evaluations of a positive-semidefinite kernel without explicitly constructing the feature vectors. In the corresponding reproducing-kernel Hilbert space, the hard-margin problem isEquivalently, its dual isThere is no upper bound because this is the hard-margin, rather than soft-margin, problem.
The successful hard-margin fit proves that the transformed observations are separable. This creates complete separation in unpenalized logistic regression: scaling a separating coefficient vector continually raises the likelihood, so no finite maximum-likelihood estimate exists. The nearly singular observed information produces the enormous reported standard errors.
A ridge-penalized logistic regression, or equivalently a Bayesian logistic model with a proper Gaussian prior, gives finite, stable coefficients. Firth bias reduction is another standard remedy.
For word-count vector , the fitted logistic model isHolding other counts fixed, one additional occurrence of "dollar" multiplies the fitted spam odds by .
The logistic classifier isThe coefficients maximize the Bernoulli logistic-regression model likelihood over the training emails.
The decision boundary is the hyperplaneAt threshold the classifier predicts spam whenso gives the original classifier. Varying translates the boundary parallel to itself: raising shrinks the region classified as spam and generally trades fewer false positives for more false negatives.
Writing , the log-likelihood isWith 500 observations and 5000 word predictors, the augmented design matrix cannot have full column rank. There are nonzero coefficient directions that leave every unchanged, so any maximizer belongs to an affine family and is not unique. In addition, the high-dimensional features may completely separate the classes, in which case no finite maximizer exists at all.
Principal component analysis centers the word-count vectors, diagonalizes their sample covariance matrix, and projects onto the eigenvectors with the largest eigenvalues. Fitting logistic regression to principal-component scores removes exact collinearity and yields a lower-dimensional design.
Each principal component is an unsupervised linear combination of many words chosen to explain predictor variance; it is not a word selected for association with spam. The dimension can be chosen from a scree plot or cumulative explained variance, but cross-validating the downstream classification loss better targets prediction.
For fixed , differentiating under the constraint givesBecause the columns are orthogonal,The fitted value is therefore , where is the orthogonal projection onto the column space of .
The residual sum of squares is minimized by choosing this column space to be the span of the leading eigenvectors ofThus may be taken to have those orthonormal eigenvectors as columns. Their nonzero scales are immaterial because inverse scaling of the scores leaves unchanged.
For student in school , "lme1" is the random-intercept linear mixed modelwhereindependently. The estimates areThe fixed intercept is the population-average expected post-study score for an untreated reference student with PTHK zero. The random intercept is school 's deviation from that population intercept.
Treatment was assigned by school, so observations within one school are clustered data and plausibly correlated. A random intercept models shared unobserved school characteristics and prevents the effective amount of treatment-level information from being exaggerated by treating all students as independent.
The fitted school standard deviation is , so the estimated between-school variation is nonzero. Accounting for it raises the TV and SC standard errors from about in "lm1" to about in "lme1", reflecting that only 28 independently randomized schools identify those effects.
Conditionally on the school effect, the response variance is . Marginally,and two distinct students in the same school have estimated covariance .
There is no single coefficient estimate for the random effect because the 28 school deviations are modeled as draws from a mean-zero distribution. The summary reports the estimated distributional variance; individual empirical Bayes predictions can be extracted separately.
Model "lme2" replaces by a random intercept and random PTHK slope:with a fitted bivariate normal covariance matrix for . This adds a random-slope variance and an intercept-slope covariance.
The likelihood-ratio statistic isAn ordinary chi-squared reference is unreliable because the null random-slope variance is on the boundary and its correlation is unidentified there; the reported correlation of one also signals a nearly singular fit. A valid practical test is a parametric bootstrap: simulate many datasets from fitted "lme1", refit both models by maximum likelihood to each, recompute the likelihood-ratio statistic, and estimate the p-value by the fraction at least . The AIC favors "lme2" slightly, whereas its BIC is larger, so the descriptive criteria do not agree.
For a model with fitted parameters and maximized likelihood , the Akaike information criterion isBackward selection starts with the full model, deletes the single variable giving the smallest AIC when that AIC is lower than the current value, and repeats until no deletion improves it. The first deletion is "indus", which gives AIC .
Adding variables can reduce training bias and increase the maximized likelihood, but it also increases estimation variance and optimism. The penalty estimates this optimism, so minimizing AIC implements a bias-variance tradeoff aimed at expected out-of-sample Kullback–Leibler performance.
For ,
At the maximum-likelihood estimates,soThere are regression coefficients and one variance parameter. Adding twice this parameter count gives
There are three regression coefficients and residual degrees of freedom, so . The reported residual standard error uses the unbiased divisor:Hence the maximum-likelihood variance estimate is . Substituting this, , and the four fitted parameters—three coefficients plus —into the formula in part c(ii) gives the AIC.
The prior isThe software samples from the Bayesian posterior; the displayed "Estimate" values are posterior means computed from the Markov-chain Monte Carlo draws.
With unit error variance, the Gaussian likelihood has log-likelihoodFull column rank gives and the least-squares identityMultiplying the likelihood by and absorbing all terms independent of into gives
For a Laplace prior , maximizing the posterior is equivalent to minimizingso the posterior mode is the Lasso. For a Gaussian prior , the mode minimizesand is the ridge regression estimator . Gaussian conjugacy makes the posterior normal, so its mode and posterior mean coincide at this ridge estimate.
The nearly identical predictor columns create severe multicollinearity. Their sum is well identified, as shown by the excellent fitted response, but their difference is weakly identified, allowing ordinary least squares to choose huge opposite coefficients with huge standard errors. The independent priors impose the ridge penalty from part c, shrinking that unstable difference toward zero and producing the stable estimates near .
The reported 95% credible interval for isConditional on the model, prior, and observed data, its posterior probability is 0.95. A frequentist 95% confidence interval instead has 95% coverage under repeated sampling before the data are observed; it does not assign a sampling probability to the fixed parameter after observing this dataset.
Articles by others on the same topic
There are currently no matching articles.