A winemaker has tasters check grape samples before the harvest. Every sample is scored [sugar, acidity, berry size, skin thickness], each from 0 to 10, and the tasters say "ready" or "wait". The grapes are ready in a few regions of sugar and acidity only, and the tasters disagree with the truth on 10 to 20% of the samples. Every vineyard has its own regions.
Write classify(X_train, y_train, X_test) that learns from the tasted samples and returns a list with one label, "ready" or "wait", for every row of X_test.
The tests call judge_harvest(classify, n, seed). It trains your function on grape_samples(n, seed), which returns (X, y) for n samples of one vineyard, and gives it 1000 hidden samples from the same vineyard, keeping their labels to itself. grape_samples(n, seed) is available in your code for experiments, and so are grow_tree(X, y, max_depth, min_size=2), the tree of the problem "A Tree for the Phone Mast", and tree_label(tree, x), which returns the tree's label for the row x.
How this problem is scored
A model passes a test when it is right on more hidden samples than always predicting the most common training label. Its quality (0 to 100) is the share of the gap between that baseline and the rule that knows the vineyard (it knows the true regions, and is still wrong wherever the tasters were) that it closes; 100 at or above that rule. The accuracies are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.
Examples
Input: grape_samples(2, 1)
Output: ([[2.1, 1.6, 5.6, 8.7], [2.6, 4.8, 2.2, 9.3]], ["ready", "wait"])
Input: judge_harvest(classify, 100, 1)
Output: a summary such as {"correct": 707, "total": 1000, "hidden_counts": {"ready": 406, "wait": 594}}
this one passes: always "wait" would be right on only 594 samples
Constraints
100 <= len(X_train) <= 500,len(X_test) = 1000- each test must finish in well under a second in your browser: a few dozen trees on 500 samples are fine
- your predictions must not depend on the clock; if you use randomness, use a
random.Randomwith a fixed seed
Goals
- Control the complexity of a decision tree with its depth and minimum node size
- Choose those settings from the training data alone, by cross-validation
- Beat the majority baseline on hidden data and get as close as possible to the best possible rule