Problem 414189 · medium · Level 04 Non-Linear Data Structures

Will the Flight Be Delayed?

logistic regression · feature scaling · gradient descent · hidden test set · majority baseline

An airport keeps a log of past departures: the scheduled hour, the distance of the flight and the wind speed, and whether the flight left late. Most flights are on time. Build a model that predicts delays.

Write predict_delays(X_train, y_train, X_test) that learns from the logged flights (X_train[i] = [hour, distance_km, wind_kmh], y_train[i] is 1 for delayed and 0 for on time) and returns a list with one prediction, 0 or 1, for every row of X_test.

The tests call judge_delays(predict_delays, n, seed). It gives your function the log flight_log(n, seed) as training data and 1000 hidden flights from the same airport as X_test, keeps their labels to itself and counts your correct predictions. flight_log(n, seed) is available in your code: explore it with Run, and measure your model on a second log with another seed, never on the data it was trained on.

How this problem is scored

A model passes a test when it is right on more hidden flights than the rule that always predicts the most common label of the training log. Its quality (0 to 100) says how much of the gap between that baseline and the best possible rule it closes: the best possible rule knows exactly how delays arise at this airport and is still wrong on some flights, because a delay is partly luck. Your accuracy on the same hidden flights is shown next to every test. Match the reference solution's quality (the par in the header) for the third star.

Examples

Input:  flight_log(4, 1)
Output: ([[13, 1097, 24], [15, 3344, 39], [6, 1843, 15], [22, 636, 31]], [0, 1, 0, 1])

Input:  judge_delays(predict_delays, 60, 1)
Output: a summary such as {"correct": 723, "total": 1000, "hidden_counts": {"1": 348, "0": 652}}
        this one passes: always 0 would be right on only 652 flights

Constraints

  • 60 <= len(X_train) <= 400, len(X_test) = 1000; hours are whole numbers from 5 to 23, distances from 200 to 4000, wind speeds from 0 to 50
  • each test must finish in well under a second in your browser: a few hundred passes of gradient descent over the training log are fine
  • your predictions must not depend on the clock; if you use randomness, use a random.Random with a fixed seed

Goals

  • Train a classifier by gradient descent on features with very different scales
  • Beat the majority baseline on flights the model has never seen
  • Improve a linear model with a feature that captures a non-linear effect
Starting Python…