Two machines fill cereal boxes. Machine 1 is precise, machine 2 is erratic, and far fewer boxes from machine 2 were weighed. The two usual t statistics for "do the machines fill the same amount on average?" are
pooled: t = (m1 - m2) / sqrt(sp2 · (1/n1 + 1/n2)), sp2 = ((n1-1) v1 + (n2-1) v2) / (n1 + n2 - 2)
Welch: t = (m1 - m2) / sqrt(v1/n1 + v2/n2), df = (v1/n1 + v2/n2)² / ((v1/n1)²/(n1-1) + (v2/n2)²/(n2-1))
where m are the group means and v the sample variances (dividing by n - 1). With samples this large, both tests reject "same average" at the 5% level when |t| > 1.96. If both machines fill the same amount on average, a good test should give a false alarm about 5% of the time. Check that by simulation.
Write false_alarms(n1, sd1, n2, sd2, trials, seed) that follows these rules exactly:
- Create one generator
rng = random.Random(seed). - Repeat
trialstimes: draw group 1 asn1valuesrng.gauss(0, sd1), then group 2 asn2valuesrng.gauss(0, sd2)(both machines have the same true mean, 0), and compute both t statistics and the Welch degrees of freedom.
Return a dict with "pooled" and "welch", the fractions of the trials in which |t| > 1.96 for each statistic, and "welch_df", the average of the Welch degrees of freedom over the trials.
Examples
Input: n1 = 3, sd1 = 1, n2 = 4, sd2 = 2, trials = 2, seed = 1
Output: {"pooled": 0.5, "welch": 0.5, "welch_df": 4.980175370338365}
Explanation: the first trial draws 1.288, 1.449, 0.066 and then -1.529, -2.184, 0.063, -2.044.
Input: n1 = 200, sd1 = 1, n2 = 50, sd2 = 3, trials = 1000, seed = 1
Output: {"pooled": 0.25, "welch": 0.06, "welch_df": 51.87648772772201}
Explanation: the pooled test raises a false alarm in a quarter of the trials: it spreads
the small, erratic group's variance over all 250 boxes and underestimates the standard error.
Input: n1 = 200, sd1 = 3, n2 = 50, sd2 = 1, trials = 1000, seed = 5
Output: {"pooled": 0.003, "welch": 0.052, "welch_df": 228.82890769456154}
Explanation: the other way round, the pooled test almost never rejects, so it would
also miss real differences.
Constraints
2 <= n1, n2,sd1, sd2 > 0,1 <= trials,(n1 + n2) * trials <= 3 * 10**5- floats are compared with a tolerance of
1e-6; use no randomness other thanrng
Goals
- Compute the pooled and the unpooled (Welch) two-sample t statistics
- Compute the Welch-Satterthwaite degrees of freedom
- Measure each test's false-alarm rate by simulating data with no real difference