A plant-health app looks at photos of leaves and names the condition. A gardener checked a batch by hand: actual[i] is the true condition of photo i and predicted[i] is the app's answer.
Write confusion_matrix(actual, predicted) that returns a tuple (labels, matrix):
labelsis the sorted list of every label that appears inactualor inpredicted;matrix[r][c]is the number of photos whose actual label islabels[r]and whose predicted label islabels[c]. Rows are actual labels, columns predicted labels.
The setup provides leaf_photos(n, seed, skill=0.7), which returns (actual, predicted) for n random photos, with conditions "healthy", "blight", "rust" and "mildew".
Examples
Input: actual = ["rust", "healthy", "rust", "blight", "healthy"]
predicted = ["blight", "healthy", "rust", "blight", "rust"]
Output: (["blight", "healthy", "rust"], [[1, 0, 0], [0, 1, 1], [1, 0, 1]])
Explanation: row "rust" has one photo called "blight" and one called "rust";
row "healthy" has one right and one called "rust". The diagonal holds the 3 correct answers.
Input: actual = ["a", "a"], predicted = ["b", "b"]
Output: (["a", "b"], [[0, 2], [0, 0]])
Explanation: "b" never occurs as a true label, but it still gets a row (of zeros) and a column.
Constraints
0 <= len(actual) == len(predicted) <= 10**5- labels are non-empty strings; there are at most 20 different labels
Goals
- Count every combination of actual and predicted label in a confusion matrix
- Lay out the matrix with actual labels as rows and predicted labels as columns
- Include labels that occur only among the predictions