A letting agency predicts the monthly rent of a flat with a linear model: for a flat with features x = [x_1, ..., x_d] (floor area, number of rooms, ...) the prediction is w_1·x_1 + ... + w_d·x_d + b. The loss of the parameters (w, b) on the training flats is the mean squared error: the average over all flats of (prediction - rent)².
Write mse_and_gradient(X, y, w, b) that returns a tuple (loss, grad_w, grad_b):
lossis the mean squared error on the rows ofXwith true valuesy;grad_wis a list with the exact partial derivative of the loss with respect to each weightw_j;grad_bis the exact partial derivative with respect tob.
Examples
Input: X = [[50, 2], [80, 3], [65, 2], [120, 4]], y = [700, 1000, 820, 1500], w = [10, 50], b = 100
Output: (850.0, [2975.0, 105.0], 40.0)
Explanation: the predictions are 700, 1050, 850 and 1500, so the errors are 0, 50, 30 and 0 and
the loss is (2500 + 900) / 4 = 850. The flats that are over-predicted pull every parameter down:
raising w_1 would raise the big errors further, so its slope is large and positive.
Input: X = [[1], [2], [3], [4], [5], [6]], y = [9, 11, 15, 16, 21, 24], w = [3], b = 5
Output: (0.8333333333333334, [-3.6666666666666665], -1.0)
Constraints
1 <= len(X) <= 2000,1 <= d <= 20, every row hasdnumbers andlen(w) == d- floats are compared with a tolerance of
1e-6
Goals
- Evaluate the mean squared error of a linear model with several features
- Derive and compute the exact gradient of that loss with respect to every weight and the bias
- Compute all the errors once and reuse them for the loss and every partial derivative