A statistics dashboard shows the weighted mean, variance, total weight and a title of a sample, refreshing them after every edit. The computations are slow on real data, so each should run again only when an attribute it uses has been reassigned. Write two descriptor classes that a class can use like this:
class Sample:
values = Watched()
weights = Watched()
label = Watched()
@derived
def mean(self):
return sum(v * w for v, w in zip(self.values, self.weights)) / self.total_weight
@derived
def total_weight(self):
return sum(self.weights)
Watched() is a plain stored attribute. Each instance has its own value; reading one that was never assigned raises AttributeError.
derived is used as a decorator on a method without arguments and turns it into a read-only attribute:
- The first read of
obj.meanruns the method and caches the result for that instance; later reads return the cached value without running the method. - A cached value stays valid until a watched attribute it depends on is assigned. It depends on the watched attributes read while it was computed, and on everything each derived attribute it read depends on (even when that one came from the cache). Assigning a watched attribute it does not depend on leaves it cached.
- If the method raises, nothing is cached and the exception reaches the caller.
- Assigning or deleting a derived attribute raises
AttributeError. - On the class itself (
Sample.mean), both kinds of descriptor return themselves.
Only assignments count as changes (the tests never modify a list in place). The tests build the class above, with four derived statistics that count their own computations, in setup helpers you can call with Run: sample_class(counts), script(ops, values=(2, 4, 4, 5), weights=None), on_the_class(), failing() and dashboard(n, seed). outcome(call) returns a result or the exception's name.
Examples
Input: script([("read", "variance"), ("read", "variance"), ("set", "label", "b"), ("read", "variance"), ("read", "title"), ("set", "weights", [1, 1, 1, 5]), ("read", "title"), ("read", "variance")])
Output: ([1.1875, 1.1875, None, 1.1875, "b (n=4)", None, "b (n=4)", 0.984375], [("mean", 2), ("title", 1), ("total_weight", 2), ("variance", 2)])
Explanation: the label is not used by the variance, and the weights are not used by the title.
Input: script([("set", "mean", 3), ("del", "total_weight"), ("read", "total_weight")])
Output: (["AttributeError", "AttributeError", 4], [("total_weight", 1)])
Constraints
- Derived methods take only
self; derived attributes may read other derived attributes, without cycles. - Everything is counted, nothing is timed.
Goals
- Write two descriptors that cooperate through state kept on each instance
- Cache a computed attribute per instance and discover its dependencies while it is computed
- Invalidate exactly the cached values that depend on an attribute when it is reassigned