A town council wants the average commute time of its residents. It cannot ask everybody, so it asks a random sample of n people and reports their mean. How much would that report change if a different sample had been drawn? Simulate it: the list population holds everyone's commute time.
Write sample_means(population, n, reps, seed) that follows these rules exactly:
- Create one generator
rng = random.Random(seed). - Run
repssurveys one after the other. Each survey picksnpeople with replacement, one callrng.choice(population)per person, and records the mean of thenvalues.
Return a tuple of three floats:
- the mean of the
repssurvey means, - their standard deviation (dividing by
reps), - the value
sigma / sqrt(n), wheresigmais the standard deviation of the whole population (dividing by its size).
The setup provides commute_times(size, seed), which returns a skewed population of commute times in minutes.
Examples
Input: population = [1, 2, 3, 4], n = 2, reps = 5, seed = 1
Output: (2.6, 1.1575836902790226, 0.7905694150420948)
Explanation: the ten calls to rng.choice give 2 1 | 3 1 | 4 4 | 4 4 | 2 1, so the five
survey means are 1.5, 2, 4, 4, 1.5. Their mean is 2.6 and their standard deviation 1.158.
The population has sigma = 1.118, and 1.118 / sqrt(2) = 0.791: five surveys are far
too few to see the pattern.
Input: population = [0, 0, 0, 10], n = 4, reps = 20000, seed = 7
Output: (2.485875, 2.16840678480192, 2.165063509461097)
Constraints
1 <= len(population) <= 10**4,1 <= n,1 <= reps, andn * reps <= 3 * 10**5- floats are compared with a tolerance of
1e-6 - use no randomness other than
rng
Goals
- Simulate the sampling distribution of the mean by repeating a survey many times
- Measure the spread of the sample means (the standard error)
- Compare it with the formula sigma / sqrt(n)