Problem 442343 · hard · Level 04 Non-Linear Data Structures

How Sure Is the Average Temperature?

standard error · dependent data · autocorrelation · batch means

A weather station records a daily value (a temperature anomaly, say) for n = 200 days, and reports the mean with a standard error. The formula s / sqrt(n) assumes that the days are independent, but weather is sticky: a warm day tends to be followed by another warm day. Other series swing the other way, or run in spells. For such data s / sqrt(n) can be far too small (or too large).

Write mean_se(values) that receives one series (a list of 200 floats, in time order) and returns your estimate of the standard error of its mean: the standard deviation that the mean of such a series would have over many repetitions of the same random process.

Each test calls se_trial(mean_se, kind, seed) from the setup. It generates 24 hidden series of the given kind ("sticky", "seesaw", "smooth", "weather" or "swing"), each from its own random process (its own strength of dependence, level and scale) whose true standard error of the mean it knows exactly, calls your function on each series, and reports the typical error of your answers: the root mean square of log(yours / true) over the 24 series. For experiments, temperature_series(kind, seed) returns one series of a kind (with other parameters than the tests).

Examples

Input:  se_trial(mean_se, "sticky", 1)
Output: for example {"error": 0.239, "naive": 0.817}
Explanation: s / sqrt(n) is typically off by a factor of about e^0.82 = 2.3 on these
series; a better estimator is off by a factor of about e^0.24 = 1.27.

How this problem is scored

  • Baseline: s / sqrt(n) with the sample standard deviation s, on the same series (the value "naive"). You pass a test when your typical error is no larger than the baseline's.
  • Best known: the smallest typical error reached on the same 24 series by any of about 25 estimators tried offline (several families, each with several settings).
  • Score: 100 · log(naive / yours) / log(naive / best), between 0 and 100 (100 at or below the best known).

Constraints

  • return a positive finite float; the judge calls your function 24 times per test, so keep each call well under 0.05 s
  • if you use randomness, use your own seeded random.Random; the result must not depend on the clock

Goals

  • See why s / sqrt(n) is wrong when neighbouring observations are related
  • Estimate the standard error of a mean from a single series of dependent values
  • Judge an estimator by its typical error over many series
Starting Python…