Problem 422650 · hard · Level 04 Non-Linear Data Structures

Predicting the Harvest

linear regression · gradient descent · feature scaling · feature engineering · hidden test set · mean squared error

An agricultural research station has run field trials: for every field it recorded the rainfall over the season, the amount of fertiliser and the mean temperature, and measured the yield. Build a model that predicts the yield of new fields.

Write predict_yield(X_train, y_train, X_test) that learns from the trials (X_train[i] = [rain_mm, fertiliser_kg_per_ha, temperature_c], y_train[i] is the yield in tonnes per hectare) and returns a list with one predicted yield (a float) for every row of X_test.

The tests call judge_yield(predict_yield, n, seed). It gives your function the trials field_trials(n, seed) as training data and 1000 hidden fields from the same region as X_test, keeps their yields to itself and measures the mean squared error of your predictions. field_trials(n, seed) is available in your code: explore it with Run, and measure your model on a second set of trials with another seed.

How this problem is scored

A model passes a test when its mean squared error on the hidden fields is lower than that of the rule that predicts the mean training yield for every field. Its quality (0 to 100) says how much of the gap between that baseline and the best possible model it closes, linearly in the error: the best possible model knows exactly how yield depends on the three measurements, and still has an error of about 0.2, because weather and soil vary in ways the measurements do not capture. The errors are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  field_trials(4, 1)
Output: ([[370, 72, 12.5], [1089, 125, 15.0], [663, 180, 32.0], [800, 27, 13.5]], [4.01, 5.03, 7.64, 4.22])

Input:  judge_yield(predict_yield, 40, 1)
Output: a summary such as {"mse": 1.065, "total": 1000, "hidden_mean": ..., "hidden_mean_sq": ...}
        this one passes: predicting the mean yield everywhere has an error of 2.05

Constraints

  • 40 <= len(X_train) <= 400, len(X_test) = 1000
  • each test must finish in well under a second in your browser: a few hundred passes of gradient descent over the trials are fine
  • your predictions must not depend on the clock; if you use randomness, use a random.Random with a fixed seed

Goals

  • Train a regression model by gradient descent on features with different scales
  • Beat the predict-the-mean baseline on fields the model has never seen
  • Discover non-linear effects in the data and let a linear model use them through new features
Starting Python…