A community garden gave three liquid feeds to tomato seedlings; groups[g] holds the heights (cm) after a month of the plants that received feed g. Not every group has the same number of plants. Do the feeds differ?
With k groups and N plants in total, let grand be the mean of all plants and mean_g the mean of group g:
SS_between = sum over groups of len(group) · (mean_g - grand)²
SS_within = sum over all plants of (height - mean of its group)²
F = (SS_between / (k - 1)) / (SS_within / (N - k))
Write anova_shuffle(groups, trials, seed) that returns a dict:
"ss_between","ss_within","f","df": the tuple(k - 1, N - k),"eta_squared":SS_between / (SS_between + SS_within), the share of all variation that lies between the groups,"p_perm": a permutation p-value. Make the listpooledof all heights (group 0 first, then group 1, and so on) and createrng = random.Random(seed). Repeattrialstimes:rng.shuffle(pooled)(the same list each time), cut it into consecutive pieces with the original group sizes in the original order, and compute F for them. Count the shuffles withF >= f - 1e-9; the p-value is(count + 1) / (trials + 1).
The setup provides plant_heights(sizes, effects, seed), which returns simulated groups: group g has sizes[g] plants whose feed adds effects[g] cm on average.
Examples
Input: groups = [[14, 17, 15], [19, 21, 18, 22], [16, 15]], trials = 9, seed = 1
Output: {"ss_between": 47.05555555555554, "ss_within": 15.166666666666668, "f": 9.307692307692303,
"df": (2, 6), "eta_squared": 0.7562499999999999, "p_perm": 0.1}
Explanation: the group means are 15.33, 20 and 15.5 around a grand mean of 17.44.
Input: groups = [[31.2, 28.5, 30.1, 33.0, 29.4], [34.8, 36.1, 33.5, 35.9], [30.5, 32.2, 29.8, 31.6, 30.9, 28.7]],
trials = 3000, seed = 2
Output: {"ss_between": 60.509500000000145, "ss_within": 24.287833333333346, "f": 14.948101587214477,
"df": (2, 12), "eta_squared": 0.7135778640837767, "p_perm": 0.0006664445184938354}
Explanation: only 1 of the 3000 shuffles produced an F this large.
Constraints
2 <= k <= 10, every group has at least 1 value,N > k, and the values within groups are not all equal1 <= trials,N * trials <= 3 * 10**5- floats are compared with a tolerance of
1e-6; use no randomness other thanrng
Goals
- Split the variation of grouped data into between-group and within-group sums of squares
- Compute the F statistic with its degrees of freedom and the share of variation explained
- Get a p-value for F by shuffling the values among groups of unequal sizes