A regional newspaper asked people in two towns whether they support a new cycle lane. In town A, yes_a of n_a people said yes; in town B, yes_b of n_b. The editor wants to know each town's support with a margin of error, and whether the towns really differ.
Write poll_intervals(yes_a, n_a, yes_b, n_b, z) that returns a dict:
"a": a tuple(p, low, high)with town A's sample proportionp = yes_a / n_aand the intervalp ± z · sqrt(p (1 - p) / n_a),"b": the same for town B,"diff": a tuple(d, low, high)withd = p_a - p_band the intervald ± z · se_d, wherese_dis the square root of the sum of the two squared standard errors (the towns were sampled independently),"verdict":"A"if the whole"diff"interval is above 0,"B"if it is entirely below 0, and"unclear"otherwise.
Do not clip the intervals to [0, 1].
Examples
Input: yes_a = 240, n_a = 400, yes_b = 180, n_b = 360, z = 1.96
Output: {"a": (0.6, 0.5519900010414497, 0.6480099989585503),
"b": (0.5, 0.44834946488391647, 0.5516505351160835),
"diff": (0.09999999999999998, 0.029482358393251834, 0.17051764160674812),
"verdict": "A"}
Input: yes_a = 52, n_a = 100, yes_b = 45, n_b = 100, z = 1.96
Output: {"a": (0.52, 0.42207843138511314, 0.6179215686148869),
"b": (0.45, 0.35249123116355124, 0.5475087688364488),
"diff": (0.07, -0.06819042513864698, 0.208190425138647),
"verdict": "unclear"}
Explanation: the difference could plausibly be anything from -6.8 to +20.8 points.
Constraints
1 <= n_a, n_b <= 10**7,0 <= yes_a <= n_a,0 <= yes_b <= n_b,z > 0- floats are compared with a tolerance of
1e-6
Goals
- Compute the standard error of a sample proportion
- Build confidence intervals for two proportions and for their difference
- Decide from an interval whether the data shows a difference