Problem 566721 · medium · Level 05 Advanced Algorithms & Graphs

Which Way Do the Intervals Miss?

confidence interval · coverage · skewed data · seeded simulation

The interval mean ± crit · s / sqrt(n) assumes roughly normal data. Waiting times, incomes and file sizes are not normal: most values are small, with a long tail of large ones. How well does the interval work for such data, and when it fails, in which direction?

Write interval_misses(dist, mu, n, crit, trials, seed) that follows these rules exactly:

  • Create one generator rng = random.Random(seed).
  • Run trials studies one after the other. Each study draws n values: rng.expovariate(1 / mu) per value if dist == "exp" (skewed, with mean mu), or rng.uniform(0, 2 * mu) per value if dist == "uniform" (symmetric, with mean mu).
  • From each study's values compute the mean m, the sample standard deviation s (dividing by n - 1) and the interval m - h to m + h with h = crit * s / sqrt(n).

An interval is below the truth when m + h < mu, above it when m - h > mu, and covers mu otherwise. Return a dict with the keys "coverage", "below" and "above" (each the fraction of the trials studies) and "width", the average width 2h of the intervals.

Examples

Input:  dist = "exp", mu = 10, n = 4, crit = 3.182, trials = 5, seed = 4
Output: {"coverage": 0.8, "below": 0.2, "above": 0.0, "width": 24.031192397471465}
Explanation: the first study draws 2.69, 1.09, 5.04, 1.68: mean 2.63, s = 1.70,
h = 3.182 * 1.70 / 2 = 2.70, so the interval ends at 5.33, below the true mean 10.

Input:  dist = "exp", mu = 4, n = 10, crit = 2.262, trials = 20000, seed = 3
Output: {"coverage": 0.9019, "below": 0.0943, "above": 0.0038, "width": 5.261285535132633}

Input:  dist = "uniform", mu = 4, n = 10, crit = 2.262, trials = 20000, seed = 3
Output: {"coverage": 0.9444, "below": 0.0277, "above": 0.0279, "width": 3.2518878241348914}
Explanation: 2.262 is the right t multiplier for 10 values. For symmetric data the 5.6%
of misses split evenly; for skewed data nearly all of them fall short of the truth.

Constraints

  • dist is "exp" or "uniform", mu > 0, 2 <= n, 1 <= trials, n * trials <= 3 * 10**5
  • floats are compared with a tolerance of 1e-6; use no randomness other than rng

Goals

  • Check the coverage of a confidence interval method by simulation
  • Split the misses into intervals that fall short of the truth and intervals that overshoot
  • See how skewed data makes the t interval miss mostly on one side
Starting Python…