Problem 484752 · medium · Level 04 Non-Linear Data Structures

Asking Half the Village

sampling without replacement · standard error · finite population correction · seeded simulation

A village council surveys households about their monthly water use. The village is small: population lists the value of every household, and each survey visits n different households (nobody is asked twice). The formula sigma / sqrt(n) for the standard error assumes independent draws, as with replacement. How does it do when the sample is a large part of the population?

Write village_survey(population, n, reps, seed) that returns a tuple of three floats:

  1. plain: sigma / sqrt(n), where sigma is the standard deviation of the population (dividing by its size N),
  2. corrected: plain · sqrt((N - n) / (N - 1)), the standard error with the finite population correction (0.0 when N = 1),
  3. simulated: the standard deviation (dividing by reps) of the means of reps simulated surveys. Create one generator rng = random.Random(seed), and draw each survey with one call rng.sample(population, n).

Examples

Input:  population = [2, 4, 6, 8], n = 2, reps = 6, seed = 1
Output: (1.5811388300841895, 1.2909944487358056, 1.2133516482134197)
Explanation: sigma = sqrt(5) = 2.236, so plain = 2.236 / sqrt(2) = 1.581 and
corrected = 1.581 · sqrt(2/3) = 1.291. The six samples are [4, 6], [2, 4], [2, 4],
[8, 4], [8, 2], [2, 4], with means 5, 3, 3, 6, 5, 3 and standard deviation 1.213.

Input:  population = [1, 2, ..., 100], n = 50, reps = 5000, seed = 3
Output: (4.082278775390039, 2.901149197588201, 2.9253410248817144)
Explanation: asking half the village, the simulated standard error is close to the
corrected value, far below sigma / sqrt(n).

Constraints

  • 1 <= n <= len(population) <= 5000, 1 <= reps, and n * reps <= 3 * 10**5
  • floats are compared with a tolerance of 1e-6; use no randomness other than rng

Goals

  • Simulate samples drawn without replacement from a small population
  • Compare the simulated standard error with sigma / sqrt(n)
  • See why sampling a large share of a population is more precise than the formula suggests
Starting Python…