Six volunteers measured their reaction time (milliseconds) in a simple clicking game, once before and once after a cup of coffee. before[i] and after[i] belong to the same person. People differ a lot in reaction time, but the question is whether each person got faster.
Write paired_t(before, after) that returns a dict:
"mean_diff": the mean of the differencesafter[i] - before[i],"sd_diff": their sample standard deviation (dividing byn - 1),"t_paired": the t statisticmean_diff / (sd_diff / sqrt(n)),"df": its degrees of freedom,n - 1,"t_unpaired": for comparison, the pooled two-sample t statistic that treatsbeforeandafteras two separate groups ofn:(mean(after) - mean(before)) / sqrt(sp2 · 2 / n), wheresp2is the average of the two groups' sample variances (dividing byn - 1).
The setup provides reaction_study(n, effect, seed), which returns a pair of lists (before, after) for n simulated volunteers whose times change by effect ms on average.
Examples
Input: before = [310, 285, 342, 298, 355, 270], after = [298, 280, 330, 291, 349, 262]
Output: {"mean_diff": -8.333333333333334, "sd_diff": 3.011090610836324, "t_paired": -6.779076806833006,
"df": 5, "t_unpaired": -0.442675527757945}
Explanation: every volunteer got 5 to 12 ms faster, a very consistent change (t = -6.8).
Treated as two unrelated groups, the same data looks like noise (t = -0.44), because the
differences between people (270 to 355 ms) swamp the effect.
Input: before = [12, 15, 11, 18], after = [14, 15, 14, 21]
Output: {"mean_diff": 2.0, "sd_diff": 1.4142135623730951, "t_paired": 2.82842712474619,
"df": 3, "t_unpaired": 0.8660254037844385}
Constraints
2 <= len(before) == len(after) <= 10**5- the differences are not all equal, and neither list is constant
- floats are compared with a tolerance of
1e-6
Goals
- Reduce paired measurements to one list of differences
- Compute the one-sample t statistic of the differences and its degrees of freedom
- See how much evidence is lost when the pairing is ignored