Problem 140558 · medium · Level 01 Prerequisites & Setup

Which Measurements Should Count?

weighted distance · nearest neighbour · feature scaling · centroid · units

A plant nursery wants a machine that names the species of a seedling ("basil", "mint" or "sage") from eight numbers recorded for it:

[height in mm, leaf width in mm, colour score 0-10, tray position in mm, shelf 1-6, pot number, delivery week 1-30, supplier code 0-99]

The machine gives a new seedling the species of the closest seedling in a list of labelled examples. Closeness uses a weighted squared distance, with one weight per feature:

distance(a, b) = w[0]·(a[0] - b[0])² + w[1]·(a[1] - b[1])² + ... + w[7]·(a[7] - b[7])²

(the closest example wins; on a tie the earlier one). With every weight equal to 1, the pot number, which runs up to 9999, decides everything, and the machine is no better than guessing.

Write choose_weights(rows, labels) that looks at the labelled examples and returns a list of 8 weights, each a number >= 0, not all 0. The judge then uses your weights and the same examples to name 200 new seedlings from the same nursery, which your function never sees, and counts how many it gets right.

seedling_data(n, seed) builds the example sets used by the tests and returns (rows, labels); it is available in your code.

How this problem is scored

Your weights pass if they name at least as many new seedlings correctly as the weights 1 / (max - min)² for every feature, where max and min are that feature's largest and smallest value in rows: this divides every feature by its range, so that all features count about equally. Your quality (0 to 100) says how much of the gap between that simple choice and the best weights we know you close. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  rows, labels = seedling_data(150, 1)
Output: a list of 8 weights such as [0.002, 0.2, 0.5, 0, 0, 0, 0, 0]

Constraints

  • 100 <= len(rows) <= 200; every row has the 8 features above
  • no randomness is needed; results must not depend on the clock

Goals

  • See how the units of each feature decide which vector is closest
  • Choose per-feature weights so that distance reflects what matters
  • Use class centroids to tell informative features from useless ones
Starting Python…