For a one-sided Wald test use the signed normal Wald statistic, rather than its square. The maximum-likelihood estimate is the sample mean , with exact variance , so
A sum of independent normal random variables is normal, giving and hence . Under the null hypothesis its mean is zero:
Thus its null distribution is exactly the standard normal distribution, without an asymptotic approximation. If the squared Wald statistic convention is used, has a chi-squared distribution with one degree of freedom; the signed form is needed to distinguish the two directions.
The same affine transformation of the normal distribution gives
Its variance stays one while its mean moves positively. This is the exact alternative distribution used to calculate statistical power. The square, if used instead, has a noncentral chi-squared distribution with one degree of freedom and noncentrality .
Let denote a quantile of the standard normal distribution. Reject when . The statistical power at the specified positive effect is
Equating this to and using gives . Thus, for the usual target ,
Rounding up ensures at least the target statistical power. This normal-mean sample size calculation assumes a positive integer sample size; if a requested power is at most , every positive sample size already exceeds that target for , and one should not square a negative quantile sum to impose an unnecessary lower bound.
Write , , and let be independent standard normal random variables obtained by centering and scaling the separate stage means. Then
The second statistic uses all patients, so the statistics are correlated even though the stages' new observations are independent. Their covariance is and each variance is one. This gives the exact bivariate normal distribution
Under the null hypothesis both means are zero. At replace by ; the mean vector is . In this group sequential design the full-sample statistic may be viewed as a potential statistic from the underlying sequence of outcomes, even on paths where recruitment stops.
Set . Under the null hypothesis rejection requires both and . Since has the standard normal distribution,
The bivariate normal distribution in the preceding calculation has a nonsingular covariance matrix and strictly positive statistical probability density everywhere. For every finite futility boundary and , the open rectangle has positive probability. Therefore
For an explicit expression, conditional on the full-sample statistic is under the null hypothesis, yielding
The Type I error is reduced because some otherwise rejecting paths stop for futility. Equality is approached as , but does not hold at a finite futility boundary.
Let and denote the first- and second-stage sample means, let , and write . Continuation is the selection event . The second-stage sample mean stays independent of , so . The first-stage sample mean has a truncated normal distribution. With and the upper-tail Inverse Mills ratio ,
The positive conditional selection bias after futility continuation comes from selecting unusually large first-stage outcomes. The unconditional sample mean of a fixed observations would be unbiased; that is a different sampling distribution from the one restricted to continued trials.
As , , and , so the estimator bias tends to zero. It decreases with : differentiating gives , because for a standard normal random variable. Thus
Use a Rao-Blackwell estimator after interim selection, which is exactly conditionally unbiased. The second-stage sample mean alone is unbiased conditional on continuation, but discards the earlier observations. Apply the Rao-Blackwell theorem by averaging conditional on the combined sample mean and the fact of continuation.
Set and . Before truncation, is , a distribution whose mean no longer involves the unknown . After imposing , its mean is . Since , the resulting estimator is
It uses the outcomes from both stages through their combined sample mean. By iterated expectation, , so its conditional estimator bias is zero, compared with the strictly positive estimator bias above. Its conditional variance is no larger than that of the second-stage-only estimate. This does not assert a smaller mean squared error than every biased estimator.
There are three free transition intensities: progression, death from the initial state, and death from the advanced state. In the three-state illness-death model, state 3 is an absorbing state, and there is no recovery transition.
Figure 1.
Irreversible three-state progression model with mild, severe and absorbing death states
.
Writing , and , with all three nonnegative, the transition intensity matrix is
Each diagonal entry is minus the sum of the row's outgoing transition intensities, rather than an additional unknown parameter. A continuous-time multi-state model with these rates describes the severity labels in the observations; an explicit cured state would require a richer state space.
In a time-homogeneous continuous-time Markov chain, a state's holding time is exponential with rate equal to its total outgoing transition intensity. Converting the specified times to months gives
The exit-type probability follows by dividing its transition intensity by the total exit rate. Therefore
These are starting values for numerical estimation, rather than further observations or constraints on the final fitted rates.
Let and , where time homogeneity removes dependence on . A recorded state at the next clinic visit contributes a transition probability; an exactly observed death contributes a statistical probability density, not the probability of being dead at that time. If the last recorded living state is , the mixed panel and exact-death likelihood factor after an interval is
This sums over the unobserved living state immediately before death.
Conditioning on the recorded initial states, the contribution of the three displayed patient histories is
In this irreversible illness-death model, , and , simplifying it to
The factor is essential: the death time is known exactly. Replacing the final statistical probability density by would instead model interval observation of death and give a different likelihood.
The assumptions are independent patient histories with common rates; the Markov property; constant rates over calendar/follow-up time in this model; the stated absence of recovery and absorption at death; accurate state labels and death times; and an observation/follow-up mechanism that is noninformative for the latent process given the observed history. Clinic dates are conditioned on. The displayed living endpoints contribute only the shown observations, with noninformative right censoring if they are follow-up endpoints. Initial state probabilities are omitted by conditioning on them. Progression between visits can be unobserved, which is precisely why panel-observed multi-state likelihood uses the matrix exponential rather than assuming a transition occurs at a visit.
The mean holding time from a transition intensity matrix is . Apply this to the fitted exit rates and use monotonic inversion for each confidence interval:
Thus the expected state durations are 86.96 months and 31.45 months, respectively. A confidence interval for a reciprocal rate reverses the endpoint order; the negative diagonal rates must first be converted to positive exit rates.
For the expected absorption time in an illness-death model, the time spent initially in the mild state is followed by an additional severe-state duration only if progression occurs before death. That probability is . Therefore
This is an unconditional mean including both possible paths to death, not a mean conditional on progression.
In the log-linear transition intensity model, each hazard ratio multiplies a specific off-diagonal transition intensity, comparing its post-transplant period with the pre-transplant period while holding the modeled origin state fixed. Diagonal entries must then be recalculated from row sums.
For mild-to-severe progression, the early hazard ratio suggests a 48% decrease, but its interval includes no effect. The later hazard ratio , with interval wholly below one, suggests a 98% decrease. These findings are compatible with suppression of progression by hematopoietic stem cell transplantation among those remaining in the mild state.
For mild-to-death, the early hazard ratio indicates a very large relative increase, and the later ratio still indicates an increase; both intervals are above one. For severe-to-death, the early ratio indicates increased mortality, whereas the later ratio indicates a 43% decrease; these intervals also exclude one. Early treatment toxicity and infection are plausible explanations for an immediate mortality increase. Later control of the underlying myelodysplastic syndrome is a plausible explanation for reduced advanced-state mortality and progression.
Relative increases must be interpreted alongside baseline rates. The early mild-state death rate is approximately per month, whereas the early severe-state death rate is per month. Thus the far larger mild-state hazard ratio partly reflects its much smaller starting mortality, rather than greater absolute mortality after transplantation. The later corresponding rates are and per month.
Within the question's model, delaying transplantation while the disease is mild can avoid a large immediate mortality cost, whereas the high baseline mortality in the severe state makes the later survival benefit more valuable. This provides a qualitative rationale for the stated policy. The estimates do not establish an optimal timing rule: treatment selection, changing health status, selection of survivors into the later period and the use of the previous visit's covariate value can affect the comparison. They are associations from the fitted cohort model, not automatically causal hazard ratios.
Let denote the full data and the pattern of indicators, with for an observed component and for a missing one. For each fixed pattern , split the data as . The missing at random condition is
Here parameterizes the missingness mechanism. Equivalently its conditional probability, with the observed data fixed, is constant over all possible completions of the missing data. The restriction is pattern-specific because the observed components depend on . Missing at random allows missingness to depend on observed values; missing completely at random imposes independence from all the data.
Ask the doctor whether, among patients with the same first-year drug-use status, patients who dropped out would be more or less likely to have used the drug in the second year than those who remained. The comparison is within the observed first-year categories, not just between all dropouts and all completers.
If there is no remaining relationship with second-year drug use after conditioning on first-year use, missing at random is plausible for these recorded variables. If treatment-related difficulties, new medication or deterioration lead to dropout in a way not accounted for by the recorded first-year status, missing at random may fail. A different response rate in the two first-year categories is allowed under missing at random. The unobserved second-year outcomes mean that the observed table alone cannot establish this assumption; clinical knowledge or additional follow-up information is needed.
Apply MAR standardization over a fully observed covariate. All first-year statuses are known, so the estimated probabilities of no use and use in year one are and . Within those categories, missing at random lets the observed second-year probabilities represent the corresponding dropout outcomes as well. They are estimated by and .
The law of total probability then gives
Equivalently, impute expected drug-use counts and for the two dropout groups, add these to the 25 observed second-year users, and divide by 102. This uses both the complete records and the fully observed first-year information.
Under missing completely at random, the observed second-year outcomes form a random subsample, so the direct pooled complete-case analysis estimate is
It is valid under missing completely at random because dropout no longer changes the marginal distribution of second-year use. Pooling is simpler than reporting separate conditional probabilities, but MCAR alone does not make this estimate more efficient than the estimate in the previous part. The fully observed first-year variable carries useful outcome information and can improve precision even when dropout is completely random.
To make the qualification explicit, write , , and for the constant response probability. For patients the leading variance of the pooled estimator is . The standardization estimator in the preceding part has leading variance
Indeed its first-order centered contribution is : the two terms have zero covariance, and their variances give that expression. The law of total variance therefore yields
The inequality is strict when there is dropout and first-year use predicts second-year use, as the distinct conditional rates suggest here. Thus the requested universal efficiency claim needs qualification. In the unrestricted joint binary-outcome model, maximizing the observed-data likelihood under either missing at random or missing completely at random gives the same standardization estimate : the all-patient first-year proportion and the two observed conditional second-year proportions maximize its factored outcome likelihood. Restricting the distinct missingness parameters to a common response probability affects their factor, not this estimate. The pooled value is the simple valid MCAR answer; retaining the first-year information gives the efficient MCAR answer and does not require replacing the previous estimate.
An ignorable missingness mechanism need not be absent from the data-generating process. It means that likelihood inference about the data-model parameters can omit the missingness factor, while still integrating over missing values.
Write the complete joint data statistical probability density as and the conditional missingness probability as . All values and only are observed. Under missing at random, is constant as varies with fixed. The actual observed-data likelihood is therefore
The assumed distinctness is understood as independent variation of and . Maximizing over multiplies by a factor independent of ; likelihood ratios, scores and likelihood curvature for are therefore unchanged by omitting . This proves likelihood ignorability.
For implementation of the observed-data likelihood with a missing covariate, a linear regression model for must be accompanied by an appropriate model for the distribution of . For independent individuals, write for the regression statistical probability density and for the age statistical probability density. Up to the ignorable factor,
Ignorability does not authorize discarding missing-age cases or assuming their ages have the distribution seen in the complete cases. It removes the need to model the observation mechanism for likelihood inference under the stated conditions, not the need to handle the missing covariates.
For a counting process adapted to its observed history , a predictable counting-process intensity specifies
or more generally is a local martingale. In survival analysis, let record failures and record the risk set, including subjects at their own observation time. Under independent right censoring and a common hazard function , the aggregate intensity is , where and .
Writing , the conditional increment equation is . Replacing by the observed increment gives the Nelson–Aalen estimator
where ranges over the distinct observed event times. Censored observations remove subjects from subsequent risk sets but do not produce hazard jumps.
For these data the Nelson–Aalen estimator gives
Thus the six fitted cumulative hazards sum to , the number of observed failures.
The event-count identity for Nelson–Aalen cumulative hazards follows by exchanging the finite sums. With everyone entering at time zero,
This includes each failure subject in its own risk set. A delayed-entry dataset requires a different risk indicator and is not covered by this particular identity.
Under a constant hazard survival model, . Imposing the same identity gives the events divided by exposure estimator
Its fitted cumulative hazards at the six times are , again summing to four. It is also the maximum-likelihood estimate under independent right censoring, since the parameter-dependent likelihood is . This is a sensible estimate if the exponential survival model is appropriate, but four failures give little precision. The event-count identity by itself does not validate constant hazard; the large final jump in the Nelson–Aalen estimator also reflects a risk set of one, rather than by itself proving an increasing hazard.
For a proper continuous event time, the survival function is . The assumed invertibility of the cumulative hazard function gives, for ,
Therefore
This is the cumulative hazard probability transformation; it applies also conditionally on a subject's covariates, using that subject's correctly specified cumulative hazard function.
The fitted transformed times are Cox–Snell residuals. They retain their event/right censoring indicators, so a right-censored residual represents an exponential observation known only to exceed its displayed value. Under the fitted model and independent right censoring, calculate the Kaplan–Meier estimator of residual survival and compare it with , or calculate the residual Nelson–Aalen estimator and compare its cumulative hazard with the diagonal . Systematic departures reveal model inadequacy; sparse extreme residual risk sets and parameter estimation require caution. Treating all censored residuals as observed event times would invalidate this diagnostic.
Without right censoring, the residual mean should be approximately one. With right censoring, use the modified Cox–Snell residual
For a true unit-rate exponential distribution, the memoryless property gives . Thus an event keeps its known transformed time, while a censored observation is replaced by the conditional expected event time. Under independent right censoring, iterated expectation makes the mean of these adjusted values one when the true hazards are used, and approximately one when fitted hazards are used. These mean-imputed values do not themselves have an exponential distribution, so the survival-curve diagnostic should still use the original censored residual dataset. Moreover fitting equations can force the adjusted sample mean to one, making its mean alone a weak diagnostic.
For the proposed mixture of a finite right censoring time and no right censoring, and
Consequently
When right censoring has positive probability this choice is unique; if , no correction is needed and any has the same effect. Its independence from is the content of exponential memorylessness: the expected extra lifetime after any right censoring time is one. Conditioning on an arbitrary independent right censoring time proves the same correction beyond this special two-point mixture.
For an event at from subject , with covariate vector , define the Schoenfeld function
The second term is the hazard-weighted mean covariate in the risk set just before the event. The Schoenfeld residual is this function evaluated at the fitted coefficient, . Calculate one residual vector per event, using every at-risk subject, including those who will subsequently be censored. There is no ordinary event residual assigned at a right censoring time.
The Cox partial likelihood score function is . At the true constant coefficient in a Cox proportional-hazards model, the conditional event subject is selected with weights proportional to , so each Schoenfeld function has conditional mean zero. If the coefficient varies with time, that centering changes. Plot residuals against event time or a transformation of it, smooth them, and investigate departures from zero. Scaled Schoenfeld residuals account for the risk-set covariate variance and can display departures in coefficient units; score function tests based on residual-time association provide a formal check. Risk-set composition affects unscaled residual variance, and the total residual score function can be zero by fitting even when a time trend is present.
Conditional on an event and the immediately preceding history, probabilities are proportional to the three instantaneous hazards. Put . The two zero-covariate subjects each have weight one, and the one-covariate subject has weight . The common baseline hazard cancels. Thus
Each individual with zero covariate has probability ; the first boxed probability is their combined probability. Conditioning on an event at a specified continuous time can be understood by the limiting conditional event probabilities in a short interval.
The hazard-weighted covariate mean is , so the Schoenfeld function at the true coefficient is
Multiplying by the two conditional probabilities gives
This verifies the score-centering property directly for this risk set.
When the event comes from the one-covariate subject, the Schoenfeld function for a three-person binary risk set is
Hence
These are limits where necessary. At a very negative coefficient the model assigns almost no event probability to the observed one-covariate subject, producing the largest positive discrepancy. At zero coefficient all three subjects have equal hazard functions, so the expected event covariate is and the discrepancy is . At a very positive coefficient the observed subject is predicted to have the event almost surely, so the discrepancy tends to zero. The function decreases strictly and stays positive at every finite coefficient: this one event alone favors increasing , while the other events in the complete Cox partial likelihood determine its overall estimate.
For mutually exclusive event types, let be the first event time and its type. The cause-specific hazard for cause is
It is a rate conditional on having had no event of any cause. The cumulative incidence function, also called the cumulative risk function, is the actual probability .
The total hazard function is , giving . Surviving every cause to time and then experiencing cause gives
In competing risks, is generally not , because competing events remove individuals before they can experience cause . This formula does not require assuming independent latent failure times for the different causes.
Write , with , and a finite . The competing risks model with transient surgical mortality has survival function
Integrating the disease cause-specific hazard against this survival function gives
The first term accounts for disease deaths while both causes act; the second includes survivors of that period who then face only disease mortality. The probability of eventually dying from surgery is
Taking the limit in gives , and therefore
The ultimate-death conclusion uses the positive continuing disease hazard. If , the surviving fraction would instead live indefinitely in this model; the asserted conclusion is not valid in that boundary case.
Use the Aalen–Johansen estimator of the cumulative incidence function, updating the overall Kaplan–Meier estimator for either event:
Right censoring changes subsequent risk sets, not the survival function by itself. Starting with the given estimates at , the complete calculation is
The middle event is of the competing type: it decreases overall survival while leaving the disease cumulative incidence function unchanged at that instant. The event-free intervals leave every estimate and risk set unchanged.
At the final time, both the event subject and the subject censored at that time are in the just-before risk set, so its denominator is eight. This is the usual event-before-censoring convention for recorded ties. After the event and right censoring, six subjects remain at risk. Consequently
The nine numbered source items are the inputs to this single calculation, not nine further questions.

Articles by others on the same topic (0)

There are currently no matching articles.