Train the one-hidden-layer network (tanh hidden units, one sigmoid output, average log loss) from a start that everybody can reproduce.
Write train_network(X, Y, hidden, seed, lr, epochs) that returns a tuple (loss, outputs): the average log loss on the training examples after training, and the list of the trained network's output probabilities for the examples of X, in order. Proceed exactly like this:
rng = random.Random(seed). Draw the input weights of hidden unit 0 (onerng.uniform(-1, 1)per input, in input order), then those of hidden unit 1, and so on; then one output weight per hidden unit, also withrng.uniform(-1, 1). All biases start at0.0.- Repeat
epochstimes: compute the gradient of the average log loss over all examples at the current weights (backpropagation), then move every weight and bias by-lrtimes its partial derivative.
ring_data(n, seed) makes the tests' second kind of dataset (a disc of 1s inside a ring of 0s); it is available in your code.
Examples
Input: X = [[0, 0], [0, 1], [1, 0], [1, 1]], Y = [0, 1, 1, 0], hidden = 3, seed = 0, lr = 0.5, epochs = 2000
Output: (0.0032771777411800046, [0.004469458324989402, 0.9978836632326161, 0.997883675636623, 0.004382457319989833])
Explanation: the network has learned XOR.
Input: the same XOR data, hidden = 2, seed = 7, lr = 0.5, epochs = 2000
Output: a loss of about 0.3481 and outputs of about [0.001, 0.499, 0.998, 0.501]
Explanation: from this start the two hidden units end in a poor local minimum and the network
cannot tell (0, 1) from (1, 1).
Constraints
1 <= len(X) <= 100, 1 to 4 inputs, 1 to 6 hidden units,0 < lr <= 5,0 <= epochs <= 3000- the output score stays within
±30 - floats are compared with a tolerance of
1e-6
Goals
- Initialise a network's weights from a seeded random generator in a fixed order
- Train with full-batch gradient descent and backpropagation for a number of epochs
- See that the same data and learning rate can end in a different minimum from a different start