Problem 254710 · hard · Level 02 Linear Data Structures

Forecasting the Heating Bill

regression · mean squared error · generalisation · hidden test set · best constant prediction

An energy supplier wants to forecast how much energy a house uses in a day for heating and hot water. For every house-day in its log it knows the outdoor temperature (°C) and the house's floor area (m²), X[i] = [temperature, area], and the energy used, y[i] in kWh. Colder days and bigger houses use more, but not in a simple straight line.

Write predict_heating(X_train, y_train, X_test) that learns from the log and returns a list with one predicted number (kWh) for every row of X_test.

The tests call judge_heating(predict_heating, n, seed). It gives your function the log heating_log(n, seed) as training data and 1000 hidden house-days drawn the same way as X_test, keeps their true values to itself and measures the mean squared error (MSE) of your predictions. heating_log(n, seed) is available in your code, so you can explore a log with Run and measure ideas on a second log with another seed.

How this problem is scored

A forecast passes a test when its MSE on the hidden days is lower than that of the simplest sensible forecast: the mean of the training values, predicted for every day (the best constant under squared error). Its quality (0 to 100) says how much of the way from that baseline to the true rule it gets, on a logarithmic scale (every halving of the error counts the same). The true rule is the formula the log is generated from; even it has an MSE of about 25 on the hidden days, because the data contain random day-to-day variation that no forecast can predict. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  heating_log(4, 1)
Output: ([[26.3, 123], [11.6, 184], [14.8, 214], [1.7, 117]], [14.0, 36.0, 21.5, 42.7])
Explanation: the small house on the frosty day (1.7 °C, 117 m²) uses more than the big house
on the mild day (14.8 °C, 214 m²).

Input:  judge_heating(predict_heating, 40, 1)
Output: a summary such as {"mse": 365.35, "total": 1000, "hidden_mean": ..., "hidden_var": ...}
        predicting the training mean every day would give an MSE of about 1006

Constraints

  • 40 <= len(X_train) <= 400, len(X_test) = 1000
  • temperatures from -5 to 30, areas from 50 to 250
  • each test must finish in well under a second in your browser: comparing every hidden day with every training day is fine
  • your predictions must not depend on the clock; if you use randomness, use a random.Random with a fixed seed

Goals

  • Build a numerical predictor from examples and apply it to data you cannot see the answers of
  • Beat the best constant prediction under squared error on new data
  • Notice when two features measured in different units need care
Starting Python…