Problem 542534 · medium · Level 05 Advanced Algorithms & Graphs

Checking the Experiment Every Day

type I error · power · optional stopping · z-test · A/B testing · seeded simulation

A website shows each visitor two designs side by side and records which one they click. If the designs are equally attractive, each visitor picks the new one with probability 0.5. After n visitors with heads clicks on the new design, the usual test statistic is z = (heads - n/2) / sqrt(n/4), and the test declares a difference when |z| > 1.96 (a 5% false-alarm rate for one look at the data).

The product manager checks the dashboard after every step visitors and stops the experiment as soon as |z| > 1.96. Simulate this habit and compare it with a single look at the end.

Write peeking(p, n_max, step, trials, seed) that follows these rules exactly:

  • Create one generator rng = random.Random(seed).
  • Run trials experiments one after the other. In each, simulate all n_max visitors in order, one call rng.random() per visitor, which picks the new design when the value is < p (always all n_max visitors, even after a stop, so that both habits see the same data).
  • Peeking: look after step, 2·step, ..., n_max visitors; the experiment stops at the first look with |z| > 1.96 (it then used that many visitors), otherwise it uses all n_max and declares nothing.
  • Single look: only the test after all n_max visitors.

Return a dict with "peeking" and "single_look", the fractions of experiments in which each habit declares a difference, and "avg_visitors", the average number of visitors the peeking habit used. With p = 0.5 these fractions are type I error rates; with p != 0.5 they measure power.

Examples

Input:  p = 0.5, n_max = 10, step = 5, trials = 4, seed = 1
Output: {"peeking": 0.25, "single_look": 0.0, "avg_visitors": 8.75}
Explanation: one of the four experiments had |z| > 1.96 after 5 visitors and stopped there.

Input:  p = 0.5, n_max = 400, step = 20, trials = 500, seed = 2
Output: {"peeking": 0.226, "single_look": 0.054, "avg_visitors": 337.8}
Explanation: the designs are identical, yet peeking 20 times "finds" a difference in 22.6%
of the experiments, more than four times the promised 5%.

Constraints

  • 0 <= p <= 1, 1 <= step <= n_max, n_max is a multiple of step, 1 <= trials, n_max * trials <= 5 * 10**5
  • floats are compared with a tolerance of 1e-6; use no randomness other than rng

Goals

  • Simulate a test's false-alarm rate and its power
  • Model an experimenter who looks at the data repeatedly and stops at the first significant result
  • Measure how much repeated looks inflate the type I error
Starting Python…