A spam filter scores an email with features x by z = w·x + b and says the email is spam with probability p = 1 / (1 + e^(-z)). Its log loss on one email is -log(p) if the email really is spam (label 1) and -log(1 - p) if it is not (label 0): a small penalty for a confident right answer, a huge one for a confident wrong answer. The loss on a data set is the average over the emails.
Write average_log_loss(X, y, w, b) that returns the average log loss of the model (w, b) on the rows of X with labels y.
The answer must be accurate even when the model is extremely sure of itself. If z = -50 for a spam email, then p is about 2·10⁻²² and the loss for it is 50; and z may be as large as ±1000, where 1 - p or e^(-z) cannot be computed directly with floats.
Examples
Input: X = [[3, 1], [0, 4], [2, 2]], y = [1, 0, 1], w = [2.0, -1.5], b = 0.5
Output: 0.07073568991414707
Explanation: the scores are 5.0, -5.5 and 1.5, all on the right side;
the losses are 0.0067, 0.0041 and 0.2014, and their average is 0.0707.
Input: X = [[10]], y = [1], w = [-5], b = 0
Output: 50.0
Constraints
1 <= len(X) <= 5000,1 <= d <= 20, labels are0or1|z| <= 1000for every row- floats are compared with a tolerance of
1e-6; do not clip probabilities, the exact loss is expected
Goals
- Compute the average log loss of a logistic model from its weights
- Evaluate the loss exactly even for extremely confident predictions
- Connect the log loss with the probability the model gives to the observed labels