The results of a hill race are in: data holds the finishing times in minutes. The organisers want to publish a "typical time" and wonder which summary would change least if a different set of runners had turned up: the mean, the median, or the trimmed mean (sort the values, drop the trim smallest and the trim largest, and average the rest).
There is only one race, so estimate each summary's standard error by resampling the data. Write steadiest_summary(data, trim, reps, seed) that follows these rules exactly:
- Create one generator
rng = random.Random(seed). - Make
repsresamples one after the other. Each resample haslen(data)values, drawn with replacement by callingrng.choice(data)once per value. - Compute all three summaries of every resample (the same resamples serve all three).
- The standard error of a summary is the standard deviation of its
repsresampled values, dividing byreps.
Return a dict with the keys "mean", "median" and "trimmed", each mapping to a tuple (value on the original data, standard error), and the key "steadiest" naming the summary with the smallest standard error (if two are within 1e-12 of each other, the one that comes first in the order mean, median, trimmed).
The median of an even number of values is the mean of the two middle ones. The setup provides race_times(n, seed), which returns the finishing times of n runners.
Examples
Input: data = [4, 8, 15, 16, 23, 42], trim = 1, reps = 5, seed = 2
Output: {"mean": (18.0, 4.524010020619612), "median": (15.5, 5.182663407939976),
"trimmed": (15.5, 5.299528280894442), "steadiest": "mean"}
Explanation: the 30 calls to rng.choice give the resamples
[4, 4, 4, 15, 8, 42], [42, 15, 15, 23, 8, 23], [4, 23, 42, 8, 16, 42],
[16, 42, 23, 15, 23, 16] and [23, 15, 4, 4, 15, 16]. Their medians are 6, 19, 19.5,
19.5 and 15, with standard deviation 5.18. Five resamples are far too few in practice.
Input: data = [3, 0, 5, 2, 0, 4, 1, 0, 3, 27, 2, 0, 6, 3, 1], trim = 2, reps = 2000, seed = 1
Output: {"mean": (3.8, 1.6404897558012634), "median": (2, 0.8509706222896299),
"trimmed": (2.1818181818181817, 1.004788436265512), "steadiest": "median"}
Explanation: the single 27 makes the mean jump around from resample to resample.
Constraints
1 <= len(data) <= 300,0 <= trimand2 * trim < len(data)1 <= reps, andlen(data) * reps <= 3 * 10**5- floats are compared with a tolerance of
1e-6; use no randomness other thanrng
Goals
- Resample a data set with replacement using a seeded generator
- Estimate the standard error of several statistics from the same resamples
- Compare how precisely the mean, the median and a trimmed mean are estimated