A bakery logged the temperature of the dough (°C) for every batch and whether the batch rose well (True) or not (False). It wants a rule of the form "if the temperature is above t, predict above, otherwise predict not above", with the fewest wrong predictions on the logged batches.
Only these thresholds are considered: for every two neighbouring distinct temperatures a < b in the log (no temperature lies strictly between them), the midpoint (a + b) / 2.
Write learn_cutoff(temps, rose) that returns a tuple (t, above, mistakes): the threshold, the prediction above (True or False) for temperatures above it, and the number of logged batches the rule gets wrong. Among equally good rules choose the smallest t, and for the same t prefer above = True.
Examples
Input: temps = [18, 22, 25, 21, 27, 24, 19], rose = [False, True, True, False, True, False, False]
Output: (21.5, True, 1)
Explanation: sorted, the log is 18 F, 19 F, 21 F, 22 T, 24 F, 25 T, 27 T. "Above 21.5 means True"
gets only the batch at 24 wrong. "Above 24.5 means True" also makes one mistake (the batch at 22),
but 21.5 is smaller.
Input: temps = [1, 2, 3, 4], rose = [True, True, False, False]
Output: (2.5, False, 0)
Constraints
2 <= len(temps) <= 500, and there are at least two distinct temperatures- temperatures are whole numbers or numbers with one decimal
Goals
- Learn a one-number rule from labelled examples by trying every useful threshold
- Try both directions of the rule
- Choose among equally good rules with a stated tie rule