Problem 323969 · medium · Level 03 Linear Management & Searching

Straighten the Curve

residuals · transformations · least-squares line · prediction error

A lab measures one quantity against another, for example the height of a seedling against the days since sowing, or a reaction rate against the concentration. The relationship is smooth but usually not a straight line: it may speed up, level off, grow like a power or like a logarithm, or bend like a parabola. Each measurement also carries some random noise.

Write predict(train_x, train_y, new_x) that learns from the 60 measurements (train_x[i], train_y[i]) and returns a list of predicted y values, one for each value in new_x (200 of them), in the same order. Your predictions are compared with the real (noisy) measurements at those points, which you never see.

All x values are positive (between about 0.5 and 40), and new_x lies within the range of train_x. The setup provides curve_trial(seed), which returns the tuple (train_x, train_y, new_x) of one experiment; each test calls predict(*curve_trial(seed)) with its own seed.

Examples

Input:  train_x = [1, 2, 3, 4, 5], train_y = [1, 4, 9, 16, 25], new_x = [6]
Output: for example [36.0]
Explanation: the straight line through these points, y = -7 + 6x, predicts 29 at x = 6, too low,
because the points curve upwards. A line fitted against x * x predicts 36.

How this problem is scored

  • For each test the checker computes the mean squared error of your predictions against the hidden measurements.
  • Baseline: the least-squares straight line y = a + b·x fitted to the 60 training points. You pass a test when your mean squared error is no larger than the baseline's.
  • Best known: the mean squared error of the true curve itself (the noise that nobody can predict).
  • Your score is 100 · log(baseline / yours) / log(baseline / best), between 0 and 100 (100 at or below the best known).

Constraints

  • return a list of 200 finite numbers
  • use only the training data given to predict; the result must not depend on the clock
  • each call must finish well within a second

Goals

  • Notice from the residuals that a straight line is the wrong shape
  • Fit a least-squares line to transformed data (x², √x, log x, 1/x, ...) and transform back
  • Choose between candidate models by their squared error
Starting Python…