A survey of n people finds k who say yes, and reports p̂ = k / n with a 95% interval. For a proportion there are only n + 1 possible outcomes, so the coverage of an interval method (the probability that its interval contains the true proportion p) can be computed exactly instead of simulated.
Two interval methods, both with multiplier z:
- the simple interval:
p̂ ± z · sqrt(p̂ (1 - p̂) / n), - the score interval: centre
c = (p̂ + z²/(2n)) / (1 + z²/n)and half-widthh = z / (1 + z²/n) · sqrt(p̂ (1 - p̂) / n + z²/(4n²)), soc ± h.
Write exact_coverage(n, p, z) that returns a dict {"simple": ..., "score": ...} with each method's exact coverage: the sum of the binomial probabilities P(k) = C(n, k) p^k (1 - p)^(n - k) over the outcomes k = 0, ..., n whose interval satisfies low <= p <= high.
Examples
Input: n = 3, p = 0.5, z = 1.96
Output: {"simple": 0.7500000000000003, "score": 1.0000000000000004}
Explanation: k = 0 and k = 3 give p̂ = 0 or 1 and a simple interval of width 0, which
misses 0.5; they have probability 1/8 each. k = 1 and k = 2 cover. The score interval
for k = 0 runs from 0 to 0.56 and covers 0.5.
Input: n = 20, p = 0.1, z = 1.96
Output: {"simple": 0.876037256000465, "score": 0.9568255047155371}
Explanation: a "95%" simple interval covers only 87.6% of the time here.
Constraints
1 <= n <= 10**5,0 < p < 1,z > 0- floats are compared with a tolerance of
1e-6 C(n, k)can be astronomically large whilep^kunderflows; compute eachP(k)in a way that stays accurate for largen
Goals
- Compute the coverage of an interval method exactly, without simulation
- Add up binomial probabilities over the outcomes whose interval contains the truth
- Compare the textbook interval for a proportion with a better one