Problem 592338 · medium · Level 05 Advanced Algorithms & Graphs

Three Feeds for the Tomato Plants

one-way ANOVA · F statistic · sums of squares · permutation test · effect size

A community garden gave three liquid feeds to tomato seedlings; groups[g] holds the heights (cm) after a month of the plants that received feed g. Not every group has the same number of plants. Do the feeds differ?

With k groups and N plants in total, let grand be the mean of all plants and mean_g the mean of group g:

SS_between = sum over groups of  len(group) · (mean_g - grand)²
SS_within  = sum over all plants of  (height - mean of its group)²
F = (SS_between / (k - 1)) / (SS_within / (N - k))

Write anova_shuffle(groups, trials, seed) that returns a dict:

  • "ss_between", "ss_within", "f",
  • "df": the tuple (k - 1, N - k),
  • "eta_squared": SS_between / (SS_between + SS_within), the share of all variation that lies between the groups,
  • "p_perm": a permutation p-value. Make the list pooled of all heights (group 0 first, then group 1, and so on) and create rng = random.Random(seed). Repeat trials times: rng.shuffle(pooled) (the same list each time), cut it into consecutive pieces with the original group sizes in the original order, and compute F for them. Count the shuffles with F >= f - 1e-9; the p-value is (count + 1) / (trials + 1).

The setup provides plant_heights(sizes, effects, seed), which returns simulated groups: group g has sizes[g] plants whose feed adds effects[g] cm on average.

Examples

Input:  groups = [[14, 17, 15], [19, 21, 18, 22], [16, 15]], trials = 9, seed = 1
Output: {"ss_between": 47.05555555555554, "ss_within": 15.166666666666668, "f": 9.307692307692303,
         "df": (2, 6), "eta_squared": 0.7562499999999999, "p_perm": 0.1}
Explanation: the group means are 15.33, 20 and 15.5 around a grand mean of 17.44.

Input:  groups = [[31.2, 28.5, 30.1, 33.0, 29.4], [34.8, 36.1, 33.5, 35.9], [30.5, 32.2, 29.8, 31.6, 30.9, 28.7]],
        trials = 3000, seed = 2
Output: {"ss_between": 60.509500000000145, "ss_within": 24.287833333333346, "f": 14.948101587214477,
         "df": (2, 12), "eta_squared": 0.7135778640837767, "p_perm": 0.0006664445184938354}
Explanation: only 1 of the 3000 shuffles produced an F this large.

Constraints

  • 2 <= k <= 10, every group has at least 1 value, N > k, and the values within groups are not all equal
  • 1 <= trials, N * trials <= 3 * 10**5
  • floats are compared with a tolerance of 1e-6; use no randomness other than rng

Goals

  • Split the variation of grouped data into between-group and within-group sums of squares
  • Compute the F statistic with its degrees of freedom and the share of variation explained
  • Get a p-value for F by shuffling the values among groups of unequal sizes
Starting Python…