A shortcut connection adds the input of a network block to its output:
is the block's input, is what its layers compute from it, and is what the block passes on. It is also called a skip connection or a residual connection. Following He et al., who introduced it, a network with shortcuts is a residual network, and the same network without them is a plain network.
Shortcuts matter here: the backward pass is where the gradient reaches each block, and where it vanishes in the plain network.
Subscripts number the block, brackets an entry from 1: W₁[2, 1] is row 2, column 1 of block 1's W, and z₁[1] the first value of its z.
Shortcuts address the vanishing gradient problem: during training, gradients shrink as they pass backward through the network, so the earliest layers barely learn.
The optimizer updates each weight using its gradient, which measures how strongly, and in which direction, a change in that weight affects the loss. Backpropagation computes these gradients by sending a gradient backward from the loss to the input, through every operation in the network. A block is any consecutive group of these operations whose output has the same shape as its input, such as any layer but the last in the card above, or one sublayer in GPT-2 together with its normalization and dropout.
Number the blocks 1 to from the input, so block computes . Going backward, each block multiplies the gradient it receives by its derivative , which says how much its output changes when its input changes, and passes the result to the block below. In a plain network, the gradient that comes out of the first block is therefore a product, with one factor per block:
If every factor halves the gradient, after 10 blocks it is of its starting size, and after 20 about one millionth. This is exponential decay: the gradient shrinks by the same fraction at every block, so the early blocks are updated by almost nothing.
With shortcuts, block computes , and each factor becomes :
The 1 is the shortcut's derivative, the derivative of with respect to itself; when is a vector, it is the identity matrix. Multiplied out, the product contains the term on its own: the gradient at the last block's output, unchanged. When the are small, each factor stays close to 1 instead of close to 0, so each block changes the gradient only a little instead of shrinking it. A shortcut does not change how a block computes its weight gradients; it changes the gradient the block receives, from which they are computed.
Limits
The shapes must match. is defined only when has the shape of . Where a block changes the shape, the shortcut must change it too, usually with a projection that has its own weights, so its derivative is no longer the identity.
Gradients can still vanish or grow. Depending on , a factor can be close to 0 or larger than 1. The values passed between blocks can grow too. Each block adds the output of its to a running sum, the residual stream, which after the last block is , so its size tends to grow with the number of blocks. This is why residual networks are paired with normalization and careful initialization.
The effective depth is shorter than the drawn depth. Multiplying out the product gives one term for each path down the network: at each block, a path takes either the 1, along the shortcut, or , through . With two blocks:
A path through many blocks multiplies many factors and shrinks like the plain network's product. Veit et al. show that most of the gradient during training comes from short paths: in their network of 54 blocks, from paths through 5 to 17 of them.
References
Sebastian Raschka, Build a Large Language Model (From Scratch) (Manning, 2024). Section 4.4 builds the network the card runs, which the code exported above follows.
Kaiming He et al., Deep Residual Learning for Image Recognition (2015). The paper the method comes from.
Andreas Veit, Michael Wilber and Serge Belongie, Residual Networks Behave Like Ensembles of Relatively Shallow Networks (2016). The path view of a residual network, and which paths carry its gradient.