Problem 653915 · medium · Level 06 Heuristics & Optimization

Training the Network From a Fixed Start

neural network · backpropagation · gradient descent · weight initialisation · XOR · local minima

Train the one-hidden-layer network (tanh hidden units, one sigmoid output, average log loss) from a start that everybody can reproduce.

Write train_network(X, Y, hidden, seed, lr, epochs) that returns a tuple (loss, outputs): the average log loss on the training examples after training, and the list of the trained network's output probabilities for the examples of X, in order. Proceed exactly like this:

  1. rng = random.Random(seed). Draw the input weights of hidden unit 0 (one rng.uniform(-1, 1) per input, in input order), then those of hidden unit 1, and so on; then one output weight per hidden unit, also with rng.uniform(-1, 1). All biases start at 0.0.
  2. Repeat epochs times: compute the gradient of the average log loss over all examples at the current weights (backpropagation), then move every weight and bias by -lr times its partial derivative.

ring_data(n, seed) makes the tests' second kind of dataset (a disc of 1s inside a ring of 0s); it is available in your code.

Examples

Input:  X = [[0, 0], [0, 1], [1, 0], [1, 1]], Y = [0, 1, 1, 0], hidden = 3, seed = 0, lr = 0.5, epochs = 2000
Output: (0.0032771777411800046, [0.004469458324989402, 0.9978836632326161, 0.997883675636623, 0.004382457319989833])
Explanation: the network has learned XOR.

Input:  the same XOR data, hidden = 2, seed = 7, lr = 0.5, epochs = 2000
Output: a loss of about 0.3481 and outputs of about [0.001, 0.499, 0.998, 0.501]
Explanation: from this start the two hidden units end in a poor local minimum and the network
cannot tell (0, 1) from (1, 1).

Constraints

  • 1 <= len(X) <= 100, 1 to 4 inputs, 1 to 6 hidden units, 0 < lr <= 5, 0 <= epochs <= 3000
  • the output score stays within ±30
  • floats are compared with a tolerance of 1e-6

Goals

  • Initialise a network's weights from a seeded random generator in a fixed order
  • Train with full-batch gradient descent and backpropagation for a number of epochs
  • See that the same data and learning rate can end in a different minimum from a different start
Starting Python…