The network of the previous problem (tanh hidden units, one sigmoid output) is stored as a dictionary with keys "W1" (one list of input weights per hidden unit), "b1", "W2" (one weight per hidden unit) and "b2". It is trained on examples X with labels Y (0 or 1) by gradient descent on the average log loss: -log(p) for an example with label 1 and -log(1 - p) for label 0, where p is the network's output.
Write train_step(net, X, Y, lr) that returns a tuple (loss, new_net): the average log loss at the current weights, and a new dictionary with the weights after one step of gradient descent with learning rate lr on all examples at once. Every weight moves by -lr times the partial derivative of the average loss with respect to it, all computed at the current weights. Do not change net.
Examples
Input: net = {"W1": [[1.0]], "b1": [0.0], "W2": [2.0], "b2": 0.0}, X = [[0.5]], Y = [0], lr = 0.1
Output: (1.2584433876122911, {"W1": [[0.9436978851164459]], "b1": [-0.11260422976710827],
"W2": [1.9669168436920883], "b2": -0.07159040902975482})
Explanation: h = tanh(0.5) = 0.4621 and p = σ(0.9242) = 0.7159, so the loss is -log(1 - 0.7159) = 1.2584.
The output error p - y = 0.7159 gives the step -0.1·0.7159 for b2 and -0.1·0.7159·0.4621 for W2.
Input: net = {"W1": [[0.5, -0.5], [0.25, 0.75]], "b1": [0.0, 0.0], "W2": [1.0, -1.0], "b2": 0.0},
X = [[1, 0], [0, 1]], Y = [1, 0], lr = 1.0
Output: (0.4392260318646378, {"W1": [[0.6753435711683108, -0.5984052530925025],
[0.04041765443287118, 0.8246485430556427]], "b1": [0.07693831807580827, -0.13493380251148612],
"W2": [1.1608549725149475, -1.0248676163458708], "b2": 0.09783017338692182})
Constraints
1 <= len(X) <= 2000, 1 to 8 inputs, 1 to 12 hidden units,0 < lr <= 10- the output score stays within
±30, sopis never exactly 0 or 1 - floats are compared with a tolerance of
1e-6
Goals
- Compute the average log loss of a one-hidden-layer network
- Send the output error back through the tanh units with the chain rule
- Update every weight from the gradient computed at the old weights