A chip factory measures two numbers on every chip early in production (a bias offset and a leakage, both already scaled to lie between -2 and 2). Weeks later the chip either passes the final test (label 1) or fails it (label 0). The pass region is bounded by a wavy curve, and near the curve the outcome is partly luck. Predict the outcome early.
Write classify(X_train, y_train, X_test) that learns from the tested chips and returns a list with one prediction, 0 or 1, per row of X_test.
The tests call judge_chips(classify, n, seed). It trains your function on chip_tests(n, seed) (a pair (X, y)) and gives it 1000 hidden chips made the same way, keeping their labels to itself. chip_tests(n, seed) is available in your code, so you can train on one seed and measure on another.
How this problem is scored
A model passes a test when it is right on at least as many hidden chips as the nearest class mean rule (predict the class whose average training point is closer). Its quality (0 to 100) is the share of the gap between that baseline and the rule that knows the generator (it knows the true curve, and still misses the chips near it that went the unlikely way) that it closes; 100 at or above that rule. The accuracies are shown next to every test. Match the reference solution's quality (the par in the header) for the third star.
Examples
Input: chip_tests(3, 1)
Output: ([[0.286, -0.284], [-1.176, 1.253], [0.614, -1.359]], [0, 1, 0])
Input: judge_chips(classify, 60, 1)
Output: a summary such as {"correct": 852, "total": 1000}
this one passes: the nearest class mean is right on 790 of these chips
Constraints
60 <= len(X_train) <= 300,len(X_test) = 1000, two features- each test must finish in well under a second in your browser: bound training by a number of epochs, not by the clock
- your predictions must not depend on the clock; if you use randomness, use a
random.Randomwith a fixed seed
Goals
- Train a small neural network on a problem that no straight line can solve
- Choose the number of hidden units, the learning rate and the number of epochs by experiment
- Beat a linear baseline on hidden data and approach the best possible accuracy