Problem 224079 · easy · Level 02 Linear Data Structures

Put the Flat Features on One Scale

standardisation · z-scores · feature scaling · columns

A letting agency describes each flat by a row of numbers, for example [area in m², rooms, age in years]. Before comparing flats, every column should be put on the same scale.

Write standardise(train, test):

  • For every column, compute the mean and the population standard deviation (divide by n) over the rows of train only.
  • Scale a value x of that column to (x - mean) / sd. If the column's standard deviation is 0 (all training values are equal), scale every value of that column to 0.0.
  • Return the tuple (scaled_train, scaled_test): both tables scaled with the training means and standard deviations, as lists of lists of floats, rows in their original order.

New flats in test are scaled with the old numbers, so their values may fall outside the training range.

The setup provides flats(n, seed), which returns n random rows [area, rooms, age] for experiments.

Examples

Input:  train = [[50, 2], [70, 3], [90, 4]], test = [[80, 3], [110, 2]]
Output: ([[-1.224744871391589, -1.224744871391589], [0.0, 0.0], [1.224744871391589, 1.224744871391589]],
         [[0.6123724356957945, 0.0], [2.449489742783178, -1.224744871391589]])
Explanation: column 0 has mean 70 and standard deviation 16.33; column 1 has mean 3
and standard deviation 0.8165. The test value 110 is (110 - 70) / 16.33 = 2.449.

Constraints

  • 1 <= len(train) <= 2000, 0 <= len(test) <= 2000
  • every row has the same number of columns, between 1 and 20
  • floats are compared with a tolerance of 1e-6

Goals

  • Standardise every column of a table to mean 0 and standard deviation 1
  • Learn the scaling from the training rows only and reuse it for new rows
  • Handle a column whose values are all the same
Starting Python…