An estate agent's data has features on very different scales: floor areas in the hundreds, ages in the tens, distances of a few kilometres. Plain gradient descent on such data needs a tiny learning rate and thousands of steps. The standard fix is to train on standardised features and translate the result back.
Write fit_standardised(X, y, lr, steps) that returns (w, b), a list of weights and a bias in the original units, so that the prediction for a raw row x is w[0]·x[0] + ... + b:
- For every feature column
jcompute its meanm_jand its standard deviations_j(divide byn, the number of rows), and replace each value byz = (x - m_j) / s_j. - On the standardised rows, run
stepssteps of batch gradient descent on the mean squared error of the modelu·z + c, starting from all weightsu_j = 0.0andc = 0.0. Each step computes the whole gradient first and then updates every parameter with learning ratelr. - Convert back: the model
u·z + cis the same function asw·x + bwithw_j = u_j / s_jandb = c - Σ u_j · m_j / s_j.
The helper flat_sales(n, seed) returns (X, y) for n flat sales, X[i] = [area in m², age in years, distance in km] and y[i] the price in thousands. Try it with Run, and try plain gradient descent on the raw features to see why scaling helps.
Examples
Input: X = [[1], [2], [3], [4], [5], [6]], y = [9, 11, 15, 16, 21, 24], lr = 0.1, steps = 50
Output: ([3.0285282033555925], 5.3999229286245924)
Explanation: 50 steps on the standardised distances reach the least-squares line w = 3.0286, b = 5.4.
Input: X = [[1, 10], [2, 30], [3, 20], [4, 40]], y = [5, 9, 8, 13], lr = 0.1, steps = 1
Output: ([0.46, 0.05], -0.65)
Constraints
2 <= len(X) <= 400,1 <= d <= 5features, every column has at least two different values0 <= steps <= 300,0 < lr <= 0.5- floats are compared with a tolerance of
1e-6
Goals
- Standardise every feature column to mean 0 and standard deviation 1
- Train a linear model by gradient descent on the standardised features
- Convert the trained weights and bias back to the original units