A website shows each visitor two designs side by side and records which one they click. If the designs are equally attractive, each visitor picks the new one with probability 0.5. After n visitors with heads clicks on the new design, the usual test statistic is z = (heads - n/2) / sqrt(n/4), and the test declares a difference when |z| > 1.96 (a 5% false-alarm rate for one look at the data).
The product manager checks the dashboard after every step visitors and stops the experiment as soon as |z| > 1.96. Simulate this habit and compare it with a single look at the end.
Write peeking(p, n_max, step, trials, seed) that follows these rules exactly:
- Create one generator
rng = random.Random(seed). - Run
trialsexperiments one after the other. In each, simulate alln_maxvisitors in order, one callrng.random()per visitor, which picks the new design when the value is< p(always alln_maxvisitors, even after a stop, so that both habits see the same data). - Peeking: look after
step, 2·step, ..., n_maxvisitors; the experiment stops at the first look with|z| > 1.96(it then used that many visitors), otherwise it uses alln_maxand declares nothing. - Single look: only the test after all
n_maxvisitors.
Return a dict with "peeking" and "single_look", the fractions of experiments in which each habit declares a difference, and "avg_visitors", the average number of visitors the peeking habit used. With p = 0.5 these fractions are type I error rates; with p != 0.5 they measure power.
Examples
Input: p = 0.5, n_max = 10, step = 5, trials = 4, seed = 1
Output: {"peeking": 0.25, "single_look": 0.0, "avg_visitors": 8.75}
Explanation: one of the four experiments had |z| > 1.96 after 5 visitors and stopped there.
Input: p = 0.5, n_max = 400, step = 20, trials = 500, seed = 2
Output: {"peeking": 0.226, "single_look": 0.054, "avg_visitors": 337.8}
Explanation: the designs are identical, yet peeking 20 times "finds" a difference in 22.6%
of the experiments, more than four times the promised 5%.
Constraints
0 <= p <= 1,1 <= step <= n_max,n_maxis a multiple ofstep,1 <= trials,n_max * trials <= 5 * 10**5- floats are compared with a tolerance of
1e-6; use no randomness other thanrng
Goals
- Simulate a test's false-alarm rate and its power
- Model an experimenter who looks at the data repeatedly and stops at the first significant result
- Measure how much repeated looks inflate the type I error