Problem 485950 · easy · Level 04 Non-Linear Data Structures

Which Way Is Downhill for the Rent Model?

mean squared error · analytic gradient · chain rule · linear model · least squares

A letting agency predicts the monthly rent of a flat with a linear model: for a flat with features x = [x_1, ..., x_d] (floor area, number of rooms, ...) the prediction is w_1·x_1 + ... + w_d·x_d + b. The loss of the parameters (w, b) on the training flats is the mean squared error: the average over all flats of (prediction - rent)².

Write mse_and_gradient(X, y, w, b) that returns a tuple (loss, grad_w, grad_b):

  • loss is the mean squared error on the rows of X with true values y;
  • grad_w is a list with the exact partial derivative of the loss with respect to each weight w_j;
  • grad_b is the exact partial derivative with respect to b.

Examples

Input:  X = [[50, 2], [80, 3], [65, 2], [120, 4]], y = [700, 1000, 820, 1500], w = [10, 50], b = 100
Output: (850.0, [2975.0, 105.0], 40.0)
Explanation: the predictions are 700, 1050, 850 and 1500, so the errors are 0, 50, 30 and 0 and
the loss is (2500 + 900) / 4 = 850. The flats that are over-predicted pull every parameter down:
raising w_1 would raise the big errors further, so its slope is large and positive.

Input:  X = [[1], [2], [3], [4], [5], [6]], y = [9, 11, 15, 16, 21, 24], w = [3], b = 5
Output: (0.8333333333333334, [-3.6666666666666665], -1.0)

Constraints

  • 1 <= len(X) <= 2000, 1 <= d <= 20, every row has d numbers and len(w) == d
  • floats are compared with a tolerance of 1e-6

Goals

  • Evaluate the mean squared error of a linear model with several features
  • Derive and compute the exact gradient of that loss with respect to every weight and the bias
  • Compute all the errors once and reuse them for the loss and every partial derivative
Starting Python…