group A
group B
still on the wrong side
background = what the network would guess there
Training 0 weights
epoch0
loss0.693
correct—
ready
Run
Learning rate
0.030
Too small and it crawls. Too large and it thrashes — the loss curve goes ragged and never settles.
Data
Shape 2-8-8-1
Hidden layers
8
Activation
Changing the shape builds a fresh network and starts over from random weights.
Inside the network each tile is what that unit sees, over the same map
forward a = act(W·x + b), then p = sigmoid(...)
loss −[y·log p + (1−y)·log(1−p)]
backward chain rule, layer by layer
update Adam, mini-batches of 32