Probability sampling chooses observational units by a specified random mechanism with known inclusion probabilities. This separates design-based representativeness from the convenience of obtaining particular observations. Simple random sampling, stratified sampling and cluster sampling are different designs with different precision and cost properties.
Cluster sampling chooses groups of observational units, such as production days, rather than dispersed individual units. Testing every unit in a selected group is a one-stage cluster sample. Positive within-cluster intraclass correlation reduces the information supplied by a fixed number of sampled units, although collecting clustered observations can be cheaper.
The design effect is a sampling variance under a complex design divided by the corresponding simple-random-sampling variance. For independent equal-size clusters of size with exchangeable intraclass correlation , a common mean-estimation approximation is . It expresses the loss of effective sample size from positive correlation; finite-population and unequal-cluster-size designs require their own calculation.
Stratified sampling partitions a population into strata and samples separately within each. It guarantees specified representation and can improve precision when the strata describe population variation. Equal sampling fractions give equal unit inclusion probabilities; unequal fractions require appropriate weighting for population summaries.
A fixed-size simple random sample gives every subset of the required size the same selection probability. Taking the first labels from a uniformly random permutation of labels produces this design. Restricting a larger uniform permutation to the desired label set preserves uniformity; biased modulo mappings should not replace uniform selection.
Articles by others on the same topic
There are currently no matching articles.