Problem 330548 · medium · Level 03 Linear Management & Searching

A Report Card for Every Disease

precision · recall · F1 score · macro average · confusion matrix · multi-class

The plant-health app from before is used on four conditions, and healthy leaves are far more common than the rest. Accuracy is dominated by the healthy leaves, so the developers want a report per condition.

Write class_report(actual, predicted) that returns a tuple (per_class, macro_f1, worst):

  • per_class is a dictionary from every label that occurs in actual or predicted to a tuple (precision, recall, f1) for that label treated as the positive class and all other labels as negative. A ratio whose denominator is 0 is 0.0, and F1 is 0.0 when precision and recall are both 0.
  • macro_f1 is the plain mean of the F1 values over all labels in per_class (0.0 for empty lists).
  • worst is the pair (actual label, predicted label) of different labels that occurs most often, or None if every prediction is right. Among equally frequent pairs, the smallest pair in the usual tuple order wins.

The setup provides leaf_photos(n, seed, skill=0.7), which returns (actual, predicted) for n random photos.

Examples

Input:  actual    = ["rust", "healthy", "rust", "blight", "healthy"]
        predicted = ["blight", "healthy", "rust", "blight", "rust"]
Output: ({"blight": (0.5, 1.0, 0.6666666666666666),
          "healthy": (1.0, 0.5, 0.6666666666666666),
          "rust": (0.5, 0.5, 0.5)},
         0.611111111111111, ("healthy", "rust"))
Explanation: "blight" is predicted twice and right once (precision 1/2), and the only blighted
leaf is found (recall 1). The mistakes (rust, blight) and (healthy, rust) occur once each;
("healthy", "rust") is the smaller pair.

Constraints

  • 0 <= len(actual) == len(predicted) <= 10**5, at most 20 different labels
  • floats are compared with a tolerance of 1e-6

Goals

  • Compute precision, recall and F1 for every class of a multi-class problem, one class against the rest
  • Average F1 over classes so that rare classes count as much as common ones
  • Find the most common kind of mistake
Starting Python…