Solution (source code)

= Solution

Choose candidate variables from subject-matter knowledge, measurement quality and the target question before screening on outcomes. For an adjusted-effect analysis, retain prespecified <confounders> and essential design variables even when their individual $P$ values are large. For prediction, consider reliable predictors available at the intended prediction time; do not include future information. Handle missingness transparently, using a justified <multiple imputation> strategy where appropriate rather than changing the analyzed population unpredictably as variables enter.

Encode a categorical variable with $K-1$ indicators relative to a stated reference category, and assess it as a group. Check <multicollinearity>, sparse categories and plausible interactions. The relevant information is chiefly the number and distribution of events, not merely the number of enrolled subjects: hundreds of correlated or rarely varying predictors cannot be supported by a small event count.

Compare prespecified nested models by a <likelihood-ratio test> based on the partial <log-likelihood> and the change in parameter dimension. Penalized methods such as <ridge regression> or <Lasso regression> applied to the Cox likelihood can stabilize many-variable models; choose tuning by suitable validation and account for mandatory variables. The <Akaike information criterion> or a limited, documented model-selection procedure can help balance fit and complexity. Unrestricted stepwise searches and repeated univariable screening invite unstable choices, omit joint confounding effects, and make naive post-selection intervals misleading. \b[Prefer a defensible, validated covariate set to a collection chosen only for small $P$ values.]