A lab measures one quantity against another, for example the height of a seedling against the days since sowing, or a reaction rate against the concentration. The relationship is smooth but usually not a straight line: it may speed up, level off, grow like a power or like a logarithm, or bend like a parabola. Each measurement also carries some random noise.
Write predict(train_x, train_y, new_x) that learns from the 60 measurements (train_x[i], train_y[i]) and returns a list of predicted y values, one for each value in new_x (200 of them), in the same order. Your predictions are compared with the real (noisy) measurements at those points, which you never see.
All x values are positive (between about 0.5 and 40), and new_x lies within the range of train_x. The setup provides curve_trial(seed), which returns the tuple (train_x, train_y, new_x) of one experiment; each test calls predict(*curve_trial(seed)) with its own seed.
Examples
Input: train_x = [1, 2, 3, 4, 5], train_y = [1, 4, 9, 16, 25], new_x = [6]
Output: for example [36.0]
Explanation: the straight line through these points, y = -7 + 6x, predicts 29 at x = 6, too low,
because the points curve upwards. A line fitted against x * x predicts 36.
How this problem is scored
- For each test the checker computes the mean squared error of your predictions against the hidden measurements.
- Baseline: the least-squares straight line
y = a + b·xfitted to the 60 training points. You pass a test when your mean squared error is no larger than the baseline's. - Best known: the mean squared error of the true curve itself (the noise that nobody can predict).
- Your score is
100 · log(baseline / yours) / log(baseline / best), between 0 and 100 (100 at or below the best known).
Constraints
- return a list of 200 finite numbers
- use only the training data given to
predict; the result must not depend on the clock - each call must finish well within a second
Goals
- Notice from the residuals that a straight line is the wrong shape
- Fit a least-squares line to transformed data (x², √x, log x, 1/x, ...) and transform back
- Choose between candidate models by their squared error