Problem 333526 · medium · Level 03 Linear Management & Searching

Name the Pressed Leaf

classification · k-nearest neighbours · feature scaling · hidden test set · majority baseline

A herbarium has thousands of pressed leaves and only a few labelled ones. Each leaf is described by seven whole numbers: length, width and stalk length (mm), weight (mg), the number of teeth along one edge, the day of the year it was collected, and the humidity of the press room (%). Write a classifier that names the tree: "oak", "maple" or "birch".

Write classify(X_train, y_train, X_test) that learns from the labelled leaves and returns a list with one name per row of X_test.

The tests call judge_leaves(classify, n, seed). It trains your function on leaf_press(n, seed) and gives it 1000 hidden leaves from the same herbarium as X_test, keeps their labels, and counts your correct answers. leaf_press(n, seed) is available in your code, so you can explore it with Run and measure your ideas, for example by training on one seed and testing on another.

How this problem is scored

A classifier passes a test when it is right on more hidden leaves than always predicting the most common label of the training leaves. Its quality (0 to 100) is the share of the gap between that baseline and the rule that knows the generator (the best rule for this data, which still makes mistakes because the species overlap) that it closes; 100 at or above that rule. The accuracies are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  leaf_press(3, 4)
Output: ([[106, 73, 34, 929, 8, 153, 34], [73, 48, 26, 560, 15, 173, 43], [70, 80, 3, 946, 11, 157, 29]],
         ["maple", "birch", "oak"])

Input:  judge_leaves(classify, 80, 1)
Output: a summary such as {"correct": 748, "total": 1000, "hidden_counts": {"oak": 497, "maple": 316, "birch": 187}}
        this one passes: always "oak" would be right on only 497 leaves

Constraints

  • 80 <= len(X_train) <= 300, len(X_test) = 1000
  • each test must finish in well under a second in your browser: comparing every hidden leaf with every training leaf is fine
  • your predictions must not depend on the clock; if you use randomness, use a random.Random with a fixed seed

Goals

  • Build a classifier whose features are measured in very different units
  • Beat the majority baseline on hidden data, and measure your ideas on data of your own
  • Decide which features and how many neighbours help, by experiment
Starting Python…