The plant-health app from before is used on four conditions, and healthy leaves are far more common than the rest. Accuracy is dominated by the healthy leaves, so the developers want a report per condition.
Write class_report(actual, predicted) that returns a tuple (per_class, macro_f1, worst):
per_classis a dictionary from every label that occurs inactualorpredictedto a tuple(precision, recall, f1)for that label treated as the positive class and all other labels as negative. A ratio whose denominator is 0 is0.0, and F1 is0.0when precision and recall are both 0.macro_f1is the plain mean of the F1 values over all labels inper_class(0.0for empty lists).worstis the pair(actual label, predicted label)of different labels that occurs most often, orNoneif every prediction is right. Among equally frequent pairs, the smallest pair in the usual tuple order wins.
The setup provides leaf_photos(n, seed, skill=0.7), which returns (actual, predicted) for n random photos.
Examples
Input: actual = ["rust", "healthy", "rust", "blight", "healthy"]
predicted = ["blight", "healthy", "rust", "blight", "rust"]
Output: ({"blight": (0.5, 1.0, 0.6666666666666666),
"healthy": (1.0, 0.5, 0.6666666666666666),
"rust": (0.5, 0.5, 0.5)},
0.611111111111111, ("healthy", "rust"))
Explanation: "blight" is predicted twice and right once (precision 1/2), and the only blighted
leaf is found (recall 1). The mistakes (rust, blight) and (healthy, rust) occur once each;
("healthy", "rust") is the smaller pair.
Constraints
0 <= len(actual) == len(predicted) <= 10**5, at most 20 different labels- floats are compared with a tolerance of
1e-6
Goals
- Compute precision, recall and F1 for every class of a multi-class problem, one class against the rest
- Average F1 over classes so that rare classes count as much as common ones
- Find the most common kind of mistake