Losses are rounded; the sum and mean are of the exact values.
The update step uses the gradients from both places E is used, added together.
The transformer blocks and the final layer normalization leave
each token with a vector of d = 768 values. Predicting the next token means choosing one of
the V = 50,257 vocabulary tokens: a multi-class classification with V classes. The output
layer projects from d to V values, one score per class, with a linear layer without a
bias:
is the output layer's weight matrix, [d, V], with one column per vocabulary token. has
V values, one per vocabulary token, each a weighted sum of the d values of with that
token's column of as the weights. These values are the logits: scores with no fixed
range, positive or negative, higher for a token the model rates as a more likely next token.
The name comes from statistics, where the logit of a probability is its log-odds, , the inverse of the sigmoid function. Neural networks use the word for any raw scores that sigmoid or softmax turns into probabilities, although only a sigmoid's input is exactly a log-odds.
For a sequence of tokens, the output layer returns a [T, V] matrix of logits, one row per
position.
Weight tying
GPT-2 has no separate . It reuses the transpose of the token embedding matrix ,
[V, d], the matrix in which the
embedding layer looks up input ids:
The output layer adds no parameters of its own, which saves V × d ≈ 38.6M parameters, about 31%
of GPT-2 (124M). Weight tying is not universal: many recent large models keep a separate output
matrix, while smaller models, where the embedding matrix is a larger share of the parameters, often
still tie.
In training, gets a gradient from each of its two uses. Through the output layer, all V × d
weights get one, because every logit enters the softmax. Through the embedding layer, the input
tokens' weights get a second one, added to the first. The optimizer then updates all V × d
weights, each with its total gradient.
Training
Softmax turns a row of logits into V probabilities , one per vocabulary token, that sum
to 1. The loss of one prediction is the cross-entropy. It compares with , the one-hot
encoding of the target: 1 at the token that actually comes next, 0 at every other. Only the target's
term is nonzero, so the loss is the negative log of its probability:
Row predicts the token at position , so the targets are the inputs shifted by one position, as in the data pipeline. The causal mask keeps every row from seeing its target, so one forward pass yields valid predictions, and an LLM is usually trained on the mean loss over every prediction in the batch.
Inference
Generation uses only the last token's row, the one that predicts a new token. It is autoregressive: the chosen token is appended to the input and the model runs again, one token per step. Choosing the token whose logit has the highest value is greedy decoding. That token also has the highest probability after softmax, so greedy decoding can skip softmax. Sampling instead draws the next token from the probabilities, usually after temperature scaling or top-k filtering of the logits.
References
Sebastian Raschka, Build a Large Language Model (From Scratch) (Manning, 2024). Chapter 4 builds the output layer and generates text with it, chapter 5 computes the loss.
Alec Radford et al., Language Models are Unsupervised Multitask Learners (2019). The model the card runs.
Ofir Press and Lior Wolf, Using the Output Embedding to Improve Language Models (2016). One of the papers that proposed weight tying.