A team compares classifiers written by different people, and nobody agreed on an interface. Some models are objects with a predict(x) method, some only have predict_all(xs) that labels a whole list at once (it is much faster, so use it whenever a model has it), and some are plain functions f(x). Some models also have a fit(xs, ys) method that must be called on the training data before predicting.
Write leaderboard(models, train, test). train and test are pairs (xs, ys). For each model:
- call
fit(train_xs, train_ys)if the model has afitmethod; - label every test point, with one call to
predict_all(test_xs)if the model has it, otherwise withpredict(x)for each point, otherwise by calling the model itself; - compute its accuracy: the fraction of test labels it got right.
The model's name is its name attribute if it has one, else its __name__ (functions have one), else the name of its class. Return a list of (name, accuracy) pairs, best accuracy first, ties by name in alphabetical order.
The setup has models of every kind for you to try: Majority(), Threshold(cut), NearestMean() (only predict_all), Both() (has both methods, and counts its calls in single and batch), the function always_one, and the dataset helper two_groups(n, seed), which returns (xs, ys).
Examples
Input: train = ([1, 2, 3, 7, 8], [0, 0, 0, 1, 1]); test = ([0, 4, 6, 9], [0, 0, 1, 1])
leaderboard([Majority(), Threshold(5), always_one, NearestMean()], train, test)
Output: [('Threshold', 1.0), ('nearest mean', 1.0), ('always_one', 0.5), ('majority', 0.5)]
Constraints
- The test set is not empty; up to
10models and5000points per set. - A model with
predict_allmust not havepredictcalled on it.
Goals
- Accept any object that behaves like a model instead of demanding one class
- Ask what an object can do with `hasattr` and `callable`, and prefer the better interface when both exist
- Train on one set and measure accuracy on another