A statistics toolkit models continuous random variables. Each distribution knows its own mean, variance and cumulative distribution function cdf(x) (the probability of a value at most x). Everything else should work for every distribution, including ones added later, without being written again.
Write a base class Distribution that cannot be created on its own (Distribution() raises TypeError), whose subclasses must provide mean(), variance() and cdf(x), and which gives every subclass:
std(): the standard deviation;prob_between(a, b): the probability of a value betweenaandb;quantile(p): the valuexwithcdf(x) = p, for0 < p < 1, accurate to1e-9;interval(level): the central interval(quantile((1 - level) / 2), quantile((1 + level) / 2));describe():"<ClassName>: mean <m>, sd <s>"with 3 decimals.
Then write three subclasses, each raising ValueError for invalid parameters:
Uniform(a, b)witha < b: equally likely anywhere betweenaandb;Exponential(rate)withrate > 0:cdf(x) = 1 - exp(-rate * x)forx >= 0, mean1 / rate, variance1 / rate ** 2;Normal(mu, sigma)withsigma > 0:cdf(x) = (1 + erf((x - mu) / (sigma * sqrt(2)))) / 2.
The tests also make a distribution of their own, a subclass of your Distribution that defines only the three required methods, with the helper logistic(mu, s). The helper raises(fn, *args) returns the name of the exception a call raises, or None.
Examples
Input: Normal(0, 1).quantile(0.975)
Output: 1.959963984540054
Input: Uniform(2, 6).describe(), Exponential(0.5).prob_between(1, 3), raises(Distribution)
Output: ('Uniform: mean 4.000, sd 1.155', 0.38340049956420363, 'TypeError')
Constraints
- Results are compared with a tolerance of
1e-6;0.001 <= p <= 0.999. - Use
math.erfandmath.exp; no other libraries.
Goals
- Write the shared algorithm once in a base class, in terms of methods each subclass supplies
- Use abstract methods to say what every subclass must provide
- Invert a cumulative distribution function numerically