A tomato greenhouse has 30 sensors: temperatures, humidities, light and CO₂ levels in different corners, each on its own scale. Many of them measure nearly the same thing, and most have nothing to do with the harvest. The grower has recorded the sensors and the day's harvest (in kg) for only a few weeks and wants to predict the harvest of new days.
Write predict(X_train, y_train, X_test): every row holds the 30 sensor readings of one day and y_train[i] is that day's harvest. Return a list with one predicted harvest (a number) per row of X_test.
The tests call judge_greenhouse(predict, n, seed). It trains your function on greenhouse(n, seed), which returns (X, y) for n days, and gives it 1000 hidden days from the same greenhouse, keeping their harvests to itself. greenhouse(n, seed) is available in your code for experiments, and so is solve(A, b), which solves a system of linear equations A·x = b given as a list of rows and a list.
How this problem is scored
A model passes a test when its mean squared error on the hidden days is lower than that of every constant prediction, even the one that knew the mean hidden harvest in advance. Its quality (0 to 100) is 100 · log(base / mse) / log(base / best), where base is the error of that best constant and best the error of the true harvest rule that generated the data (it still has an error: the harvest varies for reasons the sensors do not see). The errors are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.
Examples
Input: greenhouse(2, 1)
Output: ([[61.97, 7.63, -25.71, 33.77, -6.0, -48.21, ...], [...]], [34.1, 71.2])
Input: judge_greenhouse(predict, 30, 1)
Output: a summary such as {"mse": 485.2, "total": 1000}
this one passes: the best constant has an error of about 564.8
Constraints
30 <= len(X_train) <= 200, 30 features,len(X_test) = 1000- each test must finish in well under a second in your browser: a few dozen fits of a 30-feature model are fine
- your predictions must not depend on the clock; if you use randomness, use a
random.Randomwith a fixed seed
Goals
- See ordinary least squares with many features and few examples fail on new data
- Control the fit with a ridge penalty on standardised features
- Choose the penalty strength by cross-validation for each training set