Let indicate a recent birthday event, indicate a social gathering, and denote subsequent household COVID-19 infection. The standard instrumental variable conditions are:
The usual consistency in causal inference, well-defined exposure, and absence of relevant interference in causal inference are also needed to interpret the result causally.
The assumptions are plausible only approximately. First, birthday timing can correlate with age, household composition, season, holidays, local epidemic phase, testing, or health-care use. Any such common cause of and violates instrumental-variable independence. Second, a birthday may alter contacts, deliveries, travel, or testing even without the intended party exposure; those pathways violate the exclusion restriction. The instrument may also be weak where restrictions or low local prevalence suppress gatherings.
Two useful assessments are:
These checks cannot prove independence or exclusion, but failures directly falsify implications of those assumptions.
Political environment may be an effect modifier of the first stage: if red-county households hold larger birthday gatherings or use fewer mitigations, a truly causal contact mechanism predicts a larger birthday-associated infection increase there than in otherwise comparable blue counties. The comparison is therefore a mechanism check for heterogeneous treatment effects.
Its interpretation requires comparable epidemic timing, baseline prevalence, demographics, urbanicity, testing, reporting, and public-health rules across the compared counties, or adequate adjustment for them. It also assumes political classification changes gathering behavior without creating a different direct birthday-to-testing or birthday-to-infection pathway. Because voting category is ecological rather than individual, interpreting the pattern as individual behavior additionally risks the ecological fallacy.
First, compare counties only after matching, weighting, or regression adjustment for baseline prevalence, calendar date, population density, age structure, household size, income, testing intensity, and public-health restrictions. Use county-clustered uncertainty to respect within-county dependence.
Second, replace the coarse red/blue split by continuous vote share and estimate a prespecified birthday-event-by-vote-share interaction. A continuous analysis retains information, permits a dose-response check, and avoids sensitivity to an arbitrary 50% cutoff. Reporting subgroup sample sizes and correcting for multiple subgroup searches would further reduce selective interpretation.
A useful comparison divides counties into periods with strict and lenient limits on private gatherings. The split is worthwhile because the policy should alter the size or frequency of birthday gatherings, providing an independent check on the proposed first-stage mechanism.
If gatherings causally raise infection risk and restrictions reduce birthday contacts, the birthday-event association should be smaller under strict restrictions and larger under lenient restrictions. A graded pattern across restriction intensity would be stronger evidence than a single binary contrast.
The analysis assumes that restriction status is not merely a proxy for local epidemic severity, testing, voluntary caution, vaccination, or other determinants of infection, and that it does not change the direct effect of birthday timing on outcome ascertainment. It also assumes comparable compliance within each policy category and no differential migration or reporting.
Assess these assumptions by balancing or adjusting for pre-policy prevalence, testing, vaccination, mobility, demographics, and calendar time; inspect infection and testing trends before policy changes; and use mobility or contact data to confirm that restrictions actually weaken the birthday-to-gathering first stage. Placebo outcomes and dates provide additional negative control outcomes.
Put
At a fixed alternative , the approximate power of the Wald test increases as its variance
decreases. For fixed total sample size , the allocation problem is therefore
with integer rounding applied after solving the continuous problem. The second-order condition is positivity of the second derivative at the stationary point.
Writing gives and , so apart from the positive factor the variance is
Hence
The unique minimum is
This is the Neyman allocation for the log relative risk of failure.
The expected number of failures is
Holding the alternative and test size fixed, constant power is equivalent to fixing . The ethical allocation problem is therefore
The Lagrange multiplier stationary equations determine the ratio; the second-order condition requires the constrained stationary point to be a local minimum.
For
the stationary equations are
Dividing them gives
The fixed-power constraint then determines the total sample size. Strict convexity after eliminating one variable supplies the second-order minimum condition.
For each substudy define the log-odds treatment effect
The normal random-effects log-likelihood, up to an additive constant, is
Its score equations give
These are the maximum-likelihood estimators rather than the unbiased sample-variance estimator. At an interior solution with , the Hessian in is negative definite, which is the required second-order condition.
Given numerical values of and , maximize the joint log-likelihood over the response rates :
This is a penalized binomial regression problem and can be solved by Newton method or another numerical optimizer. One may alternate this maximization with the closed-form updates for and from part i until convergence. The selected solution should have a negative-definite Hessian in the fitted log-odds parameters.
Use a time-homogeneous continuous-time multi-state model with states (healthy), (ill), and (dead), where is absorbing. For risk-factor indicator , let the transition intensities be
Here is the infection rate without the risk factor, is the infection hazard ratio, is the recovery rate, and is the disease-death rate. The assumption that the risk factor affects only acquisition makes and common to both groups. The infinitesimal generator is
Let denote the transition semigroup of a continuous-time Markov chain. The first person is observed in at day zero, at day seven, and at day fourteen, so the contribution is
For the second person, death is observed exactly at day six but the infection time is latent. The density of an transition at day six is
Thus the combined contribution is a function of . This illustrates how panel observations contribute transition probabilities while an exactly observed transition contributes a state probability times its transition intensity.
An illness episode ends at total rate , so its mean duration is
The probability that its terminating transition is fatal is . Therefore, in day units,
Starting healthy, the first infection time is exponential with rate . Hence
and
For each , integrate the probability of occupying the ill state:
Equivalently, this is the entry of the fundamental matrix of an absorbing continuous-time Markov chain .
There is also a direct calculation. Each episode is fatal with probability , so the expected number of episodes before death is . Each lasts on average , giving
The acquisition rates change the waiting time between episodes but, under this model, not the total time eventually spent ill. Thus both risk groups have the same estimate.
One analysis can treat a reported symptom-onset date as the exact transition time. That adds an exactly observed infection-time density to the likelihood, but assumes symptoms begin immediately at infection, every relevant episode is symptomatic, and dates are recalled and reported without error.
A more realistic analysis treats true infection as a latent transition and symptom onset as a noisy observation. A reporting-delay distribution, and possibly probabilities of asymptomatic infection and non-reporting, can be added to a Hidden Markov model. Weekly tests then interval-censor the state transition while the symptom date refines its distribution. This approach uses more information but requires an identifiable and correctly specified symptom-delay and reporting model.
The risk set at contains every individual still under observation and event-free immediately before . Since the observed times are strictly ordered increasingly, these are individuals , so its size is
At time , the observed number of events is and the exposure to the common instantaneous hazard is the risk-set size . The likelihood score for a hazard increment therefore equates observed and expected events:
Thus the Nelson–Aalen estimator is
It estimates the cumulative hazard function by adding event count divided by current exposure at every observed event time.
A martingale residual is observed minus model-expected event count:
where is the fitted individual cumulative hazard. Under an adequate model it estimates the terminal value of a counting-process martingale.
Fit a model omitting the continuous explanatory variable , plot against , and add a flexible smooth curve. A curve fluctuating around zero without structure supports omission. A monotone or curved trend indicates that event incidence still depends on , suggesting inclusion of or a nonlinear transformation of it. The residuals are highly skewed, so the smoothed trend is more informative than an assumption of Gaussian scatter.
With a common hazard and the Nelson–Aalen estimator from part a,
Interchanging the order of the finite sums gives
Each hazard increment is counted once for every individual exposed to it, exactly reproducing its event count.
For the fitted Cox proportional-hazards model,
where is the fitted baseline cumulative hazard, usually obtained with the Breslow estimator.
A residual must correspond to an event and fitted cumulative event count . The event occurred much earlier than the model expected for that individual.
A residual means that the fitted expected event count by the observed follow-up time exceeded the observed count by four. It may be a censored individual with fitted cumulative hazard four, or an individual whose event occurred only after fitted cumulative hazard five. In either case the individual remained event-free substantially longer than predicted.
For an observed event,
Its supremum is one, approached when the fitted cumulative hazard is near zero, while there is no finite theoretical lower bound.
For a censored observation,
Its supremum is zero and again there is no finite theoretical lower bound. In a finite fitted dataset the realized minima are of course finite.
The residual plot has two asymmetric clouds. Event residuals lie below the horizontal boundary , while censored residuals lie at or below . Because , fitted cumulative hazard contains the rapidly increasing multiplier ; individuals with large who remain event-free until censoring can therefore have very negative residuals. At small , events tend to lie near and censored observations near . The resulting scatter is wedge-shaped and increasingly spread toward negative values as grows, rather than homoscedastic or approximately normal.
At the month-48 analysis, the observed time from treatment start is:
For example, patient D contributes months. This is right censoring at a common calendar data cutoff but with staggered treatment entry.
At duration seven there are six at risk and two deaths, at duration thirteen there are four at risk and one death, at duration twenty-six there are three at risk and one death, and at duration twenty-seven there are two at risk and one death. The Kaplan–Meier estimator is therefore
The censoring of D at duration thirty-seven causes no multiplicative drop.
At month 24 the durations and statuses are
At duration seven, two of six patients die, giving . Patient F is censored at duration nine. At duration thirteen, C dies while C, D, and A are at risk; treating an event before censoring at a tied time gives a factor . Hence
At month 36, patients F and D are censored at treatment durations 21 and 25, after which patient A is the sole remaining member of the risk set and dies at duration 26. The corresponding Kaplan–Meier factor is , so
without needing the earlier factors.
At month 48, D remains at risk past duration 26 and is eventually censored at 37, so no event empties the risk set. Part ii gives
At month 12, patients C and D are censored at durations 6 and 1, respectively. When B dies at duration 7, only A and B remain at risk, so
At month 36, all six patients have at least twelve months of potential follow-up unless they die earlier. The only deaths by duration twelve are B and E, tied at duration seven, so
In the final data, the only deaths by treatment duration twelve are B and E. Thus
The month-24 estimate already has this numerical value, so month 24 is the first twelve-monthly analysis satisfying .
Equality is not yet knowable at month 24 because patient F has only nine months of follow-up and could still die before duration twelve. Patient F reaches twelve complete months at the end of month 27, so month 36 is the earliest scheduled analysis at which the final value at duration twelve is evaluable.
Every patient must either have an observed death or at least 48 months of follow-up. The unresolved patient is D, who starts in month 12 and completes 48 months at the end of month 59. Therefore the month-60 analysis is the earliest twelve-monthly analysis at which is evaluable.
The Kaplan–Meier estimator requires independent censoring: conditional on modeled covariates, censoring must carry no information about the future event time. Censoring at a common administrative data cutoff satisfies this condition when calendar entry time is independent of prognosis. Under that assumption, the month-24 analysis has valid administrative censoring despite unequal follow-up caused by staggered entry.
The researcher should not selectively add A's post-cutoff death to an analysis explicitly defined by the month-24 data cutoff. Doing so gives extra follow-up to a patient precisely because an event became known, creating outcome-dependent ascertainment. The defensible choices are to retain the locked month-24 analysis or update every patient's record to one common later cutoff and label it as a new analysis.
Yes, a selective update would invalidate the answer to part viii: A's censoring time would be extended because A died, while follow-up for other patients remained truncated at month 24. The resulting censoring mechanism is informative. A uniform update of every patient to the same later administrative cutoff would preserve independent censoring under the original entry-time assumption.
A competing risks model describes mutually exclusive first-event types. Once one event occurs, it prevents the other event types from being observed as that individual's first event.
The cause-specific hazard for cause is
It is the instantaneous rate of cause among individuals still free of every competing event.
The cumulative incidence function for cause is
Unlike one minus a cause-specific survivor function, it is the absolute probability of observing cause by time in the presence of all competing causes.
With cause-specific hazards and , survival free of either event through time is
The probability of remaining event-free to and then experiencing in is . Therefore
For constant hazards and ,
Its limit is the probability that occurs before .
Articles were limited to the first 100 out of 119 total.

Articles by others on the same topic (0)

There are currently no matching articles.