Two classifiers, A and B, have labelled the same test set. pred_a[i] and pred_b[i] are their predictions for example i, and actual[i] is the true label.
Write compare_models(pred_a, pred_b, actual) that returns a tuple (error_a, error_b, only_a, only_b):
error_aanderror_bare the error rates (the fraction of examples predicted wrongly) of A and B, as floats;only_ais the number of examples that A gets right and B gets wrong;only_bis the number of examples that B gets right and A gets wrong.
Examples
Input: pred_a = ["cat", "dog", "dog", "cat", "cat"]
pred_b = ["cat", "cat", "dog", "dog", "cat"]
actual = ["cat", "dog", "cat", "dog", "cat"]
Output: (0.4, 0.4, 1, 1)
Explanation: A is wrong on examples 2 and 3, B on examples 1 and 2. A alone is right
on example 1, B alone on example 3. Equal error rates, but different mistakes.
Input: pred_a = [1, 1, 0], pred_b = [1, 1, 1], actual = [1, 1, 0]
Output: (0.0, 0.3333333333333333, 1, 0)
Constraints
1 <= len(actual) <= 10**5, and all three lists have the same length- labels are strings or whole numbers
Goals
- Compute the error rate of a classifier on a test set
- Compare two models example by example, not only by their totals
- See that a difference in accuracy rests on the few examples where the models disagree