Problem 392683 · hard · Level 03 Linear Management & Searching

What Should This Flat Cost?

regression · model comparison · validation set · linear regression · k-nearest neighbours · feature scaling

A letting agency wants a rent estimate for every new flat. Each flat is described by [area in m², rooms, distance to the centre in km, floor, year built], and the agency has the rents of a few hundred flats let recently. Bigger flats cost more, central ones much more, but not in a straight line.

Write predict(X_train, y_train, X_test) that learns from the let flats and returns a list with one predicted rent (a number) per row of X_test.

The tests call judge_rents(predict, n, seed). It trains your function on flat_rents(n, seed) and gives it 1000 hidden flats from the same city, keeping their rents to itself. flat_rents(n, seed) is available in your code, so you can compare models on data of your own.

How this problem is scored

A model passes a test when its mean squared error on the hidden flats is lower than that of every constant prediction, even the one that knew the mean hidden rent in advance. Its quality (0 to 100) is 100 · log(base / mse) / log(base / best), where base is the error of that best constant and best the error of the true rent rule that generated the data (it still has an error, because rents vary for reasons the five numbers do not capture). The errors are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  flat_rents(2, 1)
Output: ([[66, 3, 5.8, 7, 1976], [94, 4, 5.2, 4, 1921]], [1059, 1435])

Input:  judge_rents(predict, 80, 1)
Output: a summary such as {"mse": 29809.3, "total": 1000}
        this one passes: the best constant prediction has an error of about 161000

Constraints

  • 80 <= len(X_train) <= 300, len(X_test) = 1000
  • each test must finish in well under a second in your browser: comparing every hidden flat with every training flat is fine
  • your predictions must not depend on the clock; if you use randomness, use a random.Random with a fixed seed

Goals

  • Compare several models honestly on data held out from training
  • Combine a fitted line with k-nearest neighbours on its residuals
  • Beat every constant prediction on hidden data, and get close to the best possible error
Starting Python…