Design effect 2026-10-07
The design effect is a sampling variance under a complex design divided by the corresponding simple-random-sampling variance. For independent equal-size clusters of size with exchangeable intraclass correlation , a common mean-estimation approximation is . It expresses the loss of effective sample size from positive correlation; finite-population and unequal-cluster-size designs require their own calculation.
Past exam of the mathematics course of the University of Cambridge 2013 iii Paper 32 2 b Solution Created 2026-10-03 Updated 2026-10-07
Under the one-test-per-day interpretation, A spreads surveillance across all five days and is operationally straightforward, but uses only five tests per week and does not guarantee morning/afternoon representation. Within each day its uniform choice avoids systematically selecting a particular batch. B uses six tests, guarantees representation of both production periods, and spreads the selected batches over the week's production, though some weekdays can receive no test. C also uses six tests and may minimize the cost of collecting samples, but concentrates them on one randomly selected day.
If batches produced on the same day share contamination risks, C's six observations are positively correlated and provide less information than six dispersed batches. A simple equal-cluster-size approximation has design effect , where is within-day intraclass correlation; its effective sample size is consequently smaller than six when . C can also miss intermittent problems occurring on other days.
I would prefer B for estimating and comparing batch contamination rates, assuming positive within-day correlation and no overriding collection-cost advantage for C. It combines slightly more testing than A with explicit production-period coverage and avoids C's concentration on a single day. A is a reasonable alternative when regular daily surveillance is the priority. All three are legitimate probability samples with the stated inclusion probabilities; the preference concerns precision and coverage, not a claim that C is intrinsically biased. Their relative efficiency is not universal without a model for production variability and costs.