Problem 488858 · easy · Level 04 Non-Linear Data Structures

A Thousand Small Surveys

sampling distribution · standard error · sample mean · seeded simulation

A town council wants the average commute time of its residents. It cannot ask everybody, so it asks a random sample of n people and reports their mean. How much would that report change if a different sample had been drawn? Simulate it: the list population holds everyone's commute time.

Write sample_means(population, n, reps, seed) that follows these rules exactly:

  • Create one generator rng = random.Random(seed).
  • Run reps surveys one after the other. Each survey picks n people with replacement, one call rng.choice(population) per person, and records the mean of the n values.

Return a tuple of three floats:

  1. the mean of the reps survey means,
  2. their standard deviation (dividing by reps),
  3. the value sigma / sqrt(n), where sigma is the standard deviation of the whole population (dividing by its size).

The setup provides commute_times(size, seed), which returns a skewed population of commute times in minutes.

Examples

Input:  population = [1, 2, 3, 4], n = 2, reps = 5, seed = 1
Output: (2.6, 1.1575836902790226, 0.7905694150420948)
Explanation: the ten calls to rng.choice give 2 1 | 3 1 | 4 4 | 4 4 | 2 1, so the five
survey means are 1.5, 2, 4, 4, 1.5. Their mean is 2.6 and their standard deviation 1.158.
The population has sigma = 1.118, and 1.118 / sqrt(2) = 0.791: five surveys are far
too few to see the pattern.

Input:  population = [0, 0, 0, 10], n = 4, reps = 20000, seed = 7
Output: (2.485875, 2.16840678480192, 2.165063509461097)

Constraints

  • 1 <= len(population) <= 10**4, 1 <= n, 1 <= reps, and n * reps <= 3 * 10**5
  • floats are compared with a tolerance of 1e-6
  • use no randomness other than rng

Goals

  • Simulate the sampling distribution of the mean by repeating a survey many times
  • Measure the spread of the sample means (the standard error)
  • Compare it with the formula sigma / sqrt(n)
Starting Python…