Problem 421405 · medium · Level 04 Non-Linear Data Structures

Did the New Layout Help?

bootstrap · two samples · difference of means · percentile interval

An online shop showed its old page layout to some visitors and a new layout to others. old and new hold the minutes each visitor spent on the site. The new layout's average is higher, but is the difference more than luck?

Write layout_effect(old, new, reps, seed) that bootstraps the difference mean(new) - mean(old), following these rules exactly:

  • Create one generator rng = random.Random(seed).
  • Repeat reps times: first resample old (len(old) calls rng.choice(old)), then resample new (len(new) calls rng.choice(new)), and record the difference of their means, new minus old.

Return a tuple (observed, se, not_better, (low, high)):

  • observed: the difference on the original data,
  • se: the standard deviation of the reps bootstrap differences (dividing by reps),
  • not_better: the fraction of bootstrap differences that are <= 0,
  • (low, high): the middle 95% of the bootstrap differences: sort them, let c = int(0.025 * reps), and take the values at positions c and reps - 1 - c.

The setup provides shop_visits(n_old, n_new, lift, seed), which returns a pair of lists (old, new) in which the new layout adds lift minutes on average.

Examples

Input:  old = [3, 1, 4, 1, 5], new = [9, 2, 6], reps = 4, seed = 1
Output: (2.866666666666667, 1.8752036926395195, 0.25, (-1.2000000000000002, 3.866666666666667))
Explanation: the four rounds draw [1, 5, 3, 4, 3] then [2, 2, 2]; [1, 1, 3, 1, 3] then
[2, 2, 6]; [3, 1, 4, 1, 5] then [9, 2, 9]; [3, 3, 5, 3, 1] then [6, 9, 2]. The differences
are -1.2, 1.53, 3.87 and 2.67; one of the four is <= 0, and with c = 0 the range is
the smallest to the largest.

Input:  old = [0, 1, 0, 0, 1, 0, 0, 0, 1, 0], new = [1, 0, 1, 1, 0, 1, 0, 1], reps = 2000, seed = 7
Output: (0.325, 0.22778018460788024, 0.1025, (-0.125, 0.775))
Explanation: 1 means the visitor bought something. The new layout converts 62.5% against
30%, but with so few visitors about one bootstrap round in ten shows no improvement.

Constraints

  • 1 <= len(old), len(new), 1 <= reps, and (len(old) + len(new)) * reps <= 4 * 10**5
  • floats are compared with a tolerance of 1e-6; use no randomness other than rng

Goals

  • Bootstrap a statistic that depends on two independent samples
  • Resample each group separately, keeping the group sizes
  • Summarise the bootstrap distribution by its spread, a tail share and a middle-95% range
Starting Python…