A water pump's vibration sensor produces one reading per minute. The pump runs in a few operating modes (idle, normal load, heavy load, ...), each with its own typical level and spread, but the log does not say which mode it was in, and nobody knows how many modes there are. The maintenance team wants a probability density for the readings, so that later they can flag readings that the model finds very unlikely.
Write fit_density(readings) that returns a density model as a list of components (weight, mean, sd): the model says a reading comes from component j with probability weight_j, and is then normal with that mean and standard deviation. Use between 1 and 10 components; the weights must be positive and add up to 1, and every sd must be at least 0.01.
Each test calls density_trial(fit_density, case) from the setup. It gives your function the training readings of test case (250 to 600 of them), then measures your model on 4000 new readings from the same pump that your function never saw: the average log density "heldout" (higher is better). You can look at the training readings with sensor_sample(case), and call density_trial yourself with Run.
Examples
Input: density_trial(fit_density, 1)
Output: for example {"heldout": -4.20, "single": -4.451, "true": -4.137, "readings": 302}
Explanation: one normal curve fitted to the training readings scores -4.451 per new
reading; the density that generated the readings scores -4.137.
How this problem is scored
- Baseline: a single normal distribution with the mean and standard deviation (dividing by n) of the training readings; its held-out score is
"single". You pass a test when your"heldout"is at least as high. - Best known: the held-out score of the true density that generated the readings (
"true"). - Score:
100 * (heldout - single) / (true - single), from 0 at the baseline to 100 at (or above) the true density.
Constraints
- 250 to 600 training readings per test
- the result must not depend on the clock; your function must finish well within a second in the browser (bound your loops by numbers of rounds and restarts)
Goals
- Fit a mixture of normal distributions to unlabelled data
- Choose the number of components without overfitting
- Judge a density model by its log-likelihood on data it has not seen