Problem 440758 · medium · Level 04 Non-Linear Data Structures

When Does the Bell Curve Fit the Bars?

binomial distribution · normal approximation · continuity correction · cumulative distribution function

Before computers, binomial probabilities with large n were read off a normal table. The normal curve used is the one with the same mean n·p and standard deviation sqrt(n·p·(1 - p)) as the binomial, and each whole number k is treated as the stretch from k - 0.5 to k + 0.5 (the continuity correction). How good is that?

Write bell_check(n, p, a, b) that returns a dict:

  • "exact": the exact binomial probability P(a <= K <= b) (values of a below 0 or b above n simply include everything on that side),
  • "approx": the normal approximation Φ(b + 0.5) - Φ(a - 0.5), where Φ is the CDF of the normal distribution with the binomial's mean and standard deviation,
  • "worst": the largest error of the approximation over the whole distribution, max |P(K <= k) - Φ(k + 0.5)| over k = 0, 1, ..., n,
  • "worst_k": the k where that largest error occurs; if several k are within 1e-9 of it, the smallest of them.

Examples

Input:  n = 10, p = 0.5, a = 4, b = 6
Output: {"exact": 0.65625, "approx": 0.6572182888520886, "worst": 0.0026861602537622264, "worst_k": 1}
Explanation: P(4 <= K <= 6) = (210 + 252 + 210) / 1024 = 0.65625. The normal curve with
mean 5 and sd 1.581 has area 0.6572 between 3.5 and 6.5. The largest error, 0.0027,
occurs at k = 1 (and, by symmetry, at k = 8).

Input:  n = 100, p = 0.03, a = 0, b = 2
Output: {"exact": 0.41977508298185423, "approx": 0.36462322789126855, "worst": 0.035054208388640706, "worst_k": 2}
Explanation: with a mean of only 3 the distribution is skewed and the bell curve is poor.

Constraints

  • 1 <= n <= 1000, 0 < p < 1
  • a <= b are integers (they may lie outside 0 .. n)
  • floats are compared with a tolerance of 1e-6

Goals

  • Approximate a binomial distribution by a normal one with the same mean and variance
  • Apply the continuity correction when a continuous curve stands in for whole numbers
  • Measure how far the approximation is from the exact distribution
Starting Python…