A shop compares the delivery times (minutes) of two couriers: a for courier A and b for courier B. Courier B looks slower, but courier A once lost a parcel for three hours. Is there a real difference? Test it without any distribution formula, by shuffling.
Write courier_test(a, b, trials, seed) that follows these rules exactly:
- The observed differences are
median(b) - median(a)andmean(b) - mean(a). The median of an even number of values is the mean of the two middle ones. - Make the list
pooled = a + band createrng = random.Random(seed). - Repeat
trialstimes: callrng.shuffle(pooled)(shuffling the same list again each time, without restoring it), treatpooled[:len(a)]as courier A's times and the rest as courier B's, and compute both differences again. - A shuffle counts as at least as extreme for a statistic when the absolute value of its difference is
>=the absolute observed difference minus1e-9. - The p-value of a statistic is
(count + 1) / (trials + 1): the observed labelling is one of the possible labellings too, so it is counted, and a permutation p-value is never 0.
Return a dict with "median_diff", "p_median", "mean_diff" and "p_mean".
The setup provides delivery_times(n_a, n_b, shift, seed), which returns a pair (a, b) of simulated delivery times in which courier B is shift minutes slower on a typical day, and about one delivery in ten is very late.
Examples
Input: a = [31, 28, 35, 30, 150, 33], b = [36, 38, 34, 41, 37], trials = 9, seed = 1
Output: {"median_diff": 5.0, "p_median": 0.1, "mean_diff": -13.966666666666661, "p_mean": 1.0}
Input: a = [29, 33, 31, 36, 30, 28, 34, 190, 32, 35, 30, 27],
b = [36, 39, 35, 41, 38, 34, 40, 37, 42, 33], trials = 5000, seed = 2
Output: {"median_diff": 6.0, "p_median": 0.004799040191961607, "mean_diff": -7.083333333333336, "p_mean": 1.0}
Explanation: courier B is typically 6 minutes slower, and the medians show it clearly.
By the means, B is faster, and every shuffle gives a difference at least as large:
the 190 dominates the mean of whichever group it lands in.
Constraints
1 <= len(a), len(b),1 <= trials,(len(a) + len(b)) * trials <= 4 * 10**5- floats are compared with a tolerance of
1e-6; use no randomness other thanrng
Goals
- Carry out a permutation test for two groups of different sizes with a seeded shuffle
- Use the same shuffles to test two different statistics
- See how a single outlier can hide a difference in means but not in medians