A small neural network has one hidden layer. It is stored as a dictionary net:
net["W1"]: one list of input weights per hidden unit (solen(net["W1"])is the number of hidden units and every inner list has one weight per input),net["b1"]: one bias per hidden unit,net["W2"]: one output weight per hidden unit,net["b2"]: the output bias (a number).
For an input x, hidden unit j computes h_j = tanh(W1[j] · x + b1[j]), and the network's output is the probability σ(W2 · h + b2) with σ(z) = 1 / (1 + e^(-z)).
Write predict_proba(net, X) that returns the list of output probabilities, one per input in X.
Examples
Input: net = {"W1": [[1.0, -1.0], [0.5, 0.5]], "b1": [0.0, -0.5], "W2": [2.0, -1.0], "b2": 0.25},
X = [[0, 0], [1, 0], [0, 1]]
Output: [0.6708688052847959, 0.8548537208454055, 0.2187119542210579]
Explanation: for [1, 0] the hidden values are tanh(1) = 0.7616 and tanh(0) = 0,
so the output is σ(2·0.7616 - 0 + 0.25) = σ(1.7732) = 0.8549.
Input: net = {"W1": [[4.0, 4.0], [4.0, 4.0]], "b1": [-2.0, -6.0], "W2": [5.0, -5.0], "b2": -4.0},
X = [[0, 0], [0, 1], [1, 0], [1, 1]]
Output: [0.02145310455528502, 0.9964606826037798, 0.9964606826037798, 0.02145310455528502]
Explanation: a hand-made network for XOR. The first hidden unit says "at least one input is 1",
the second "both are 1", and the output is high only when the first is on and the second off.
Constraints
1 <= len(X) <= 5000, 1 to 10 inputs, 1 to 20 hidden units|W2 · h + b2| <= 30for every input- floats are compared with a tolerance of
1e-6
Goals
- Compute a hidden layer of tanh units from weights and biases
- Combine the hidden values into one sigmoid output probability
- See how a hidden layer lets a network compute XOR