A letting agency describes each flat by a row of numbers, for example [area in m², rooms, age in years]. Before comparing flats, every column should be put on the same scale.
Write standardise(train, test):
- For every column, compute the mean and the population standard deviation (divide by
n) over the rows oftrainonly. - Scale a value
xof that column to(x - mean) / sd. If the column's standard deviation is 0 (all training values are equal), scale every value of that column to0.0. - Return the tuple
(scaled_train, scaled_test): both tables scaled with the training means and standard deviations, as lists of lists of floats, rows in their original order.
New flats in test are scaled with the old numbers, so their values may fall outside the training range.
The setup provides flats(n, seed), which returns n random rows [area, rooms, age] for experiments.
Examples
Input: train = [[50, 2], [70, 3], [90, 4]], test = [[80, 3], [110, 2]]
Output: ([[-1.224744871391589, -1.224744871391589], [0.0, 0.0], [1.224744871391589, 1.224744871391589]],
[[0.6123724356957945, 0.0], [2.449489742783178, -1.224744871391589]])
Explanation: column 0 has mean 70 and standard deviation 16.33; column 1 has mean 3
and standard deviation 0.8165. The test value 110 is (110 - 70) / 16.33 = 2.449.
Constraints
1 <= len(train) <= 2000,0 <= len(test) <= 2000- every row has the same number of columns, between 1 and 20
- floats are compared with a tolerance of
1e-6
Goals
- Standardise every column of a table to mean 0 and standard deviation 1
- Learn the scaling from the training rows only and reuse it for new rows
- Handle a column whose values are all the same