With one A and one B observation per stratum and no ties, the stratified log-rank statistic receives a nonzero contribution only when the earlier observation is a failure while both remain at risk. It is if A fails first and if B fails first. After the first observation, only one member remains and every further observed-minus-expected contribution is zero. Under an exchangeable null with independent censoring, conditional on being informative the two signs have equal probability, with mean zero and variance . Thus standardization is the normal approximation to a sign test on informative pairs. Failure-time distances do not affect this statistic.
Use the A-group observed-minus-expected convention for the stratified log-rank statistic. At the first failure both teeth are at risk, so A has one observed event and null expectation , giving . At the later B failure, A is no longer at risk; the A contribution is zero. This patient's contribution is , favoring B. The opposite group convention reverses the sign but gives the same test.
Only sets 4, 5, 7 and 8 have an earlier failure with both members at risk. Sets 4 and 8 have B fail first and favor A, giving informative pairs. Sets 5 and 7 have A fail first and favor B, giving . Both-censored pairs contribute zero; in sets 3 and 6 the earlier observation is censoring, so the later failure occurs with no remaining comparator and also contributes zero.
There are thus 36 informative independent strata. The stratified log-rank statistic, its null variance, and its standardized value are
Equivalently the squared statistic is nine. The normal approximation gives a two-sided p-value using the stated bound. The exact conditional sign test instead gives ; the slight difference is discreteness rather than a different direction of effect. There is strong evidence favoring Method B, which is associated with the later failure in three quarters of the informative pairs, under the test assumptions.
At each event time within a stratum, let be the group risk sets and let be the total number of observed failures. Under the null of equal hazard functions, the expected A count is . The log-rank statistic adds observed minus expected A counts over event times; the stratified log-rank statistic first makes these comparisons within each patient stratum, then adds across patients. Let be the sum of the corresponding null conditional variances. The standardized statistic has an approximate standard normal distribution under the null, and its square has an approximate one-degree-of-freedom chi-squared distribution. Compare it with the appropriate quantile for a two-sided test, under the usual independent censoring and exchangeability assumptions.