Part 8 of 83 min read

Output Layer

The linear layer that turns each token's vector into a next-token score for every token in the vocabulary.

LLMGPToutput layerlogitscross-entropyfundamentalsplayground
mode
batch size
values
r50k_base
ttokenpredictionstargetrank / 50,257loss
sequence 1 of 1input "the match did not" · targets " match did not match"
1the1169,0.014 match1,1038.89
2 match2872 was0.114 did546.09
3 did750 not0.773 not10.26
4 not407 go0.084 match134.21
batch loss4.861
mean of 1 × 4 = 4 token losses

Losses are rounded; the sum and mean are of the exact values.

gradients with weight tying: the output layer and the embedding share E [V, d] = [50257, 768]
output layerall 50,257 rows × 768 weights
embedding layer4 rows × 768 weights: the input ids
sumall 50,257 rows; the 4 marked rows get both gradients
id 0id 50,256

The update step uses the gradients from both places E is used, added together.

input ids[1, 4]
token vectors[1, 4, 768]
logits[1, 4, 50257]
targets[1, 4]
lossscalar

The transformer blocks and the final layer normalization leave each token with a vector hth_t of d = 768 values. Predicting the next token means choosing one of the V = 50,257 vocabulary tokens: a multi-class classification with V classes. The output layer projects hth_t from d to V values, one score per class, with a linear layer without a bias:

st=htWs_t = h_t W

WW is the output layer's weight matrix, [d, V], with one column per vocabulary token. sts_t has V values, one per vocabulary token, each a weighted sum of the d values of hth_t with that token's column of WW as the weights. These values are the logits: scores with no fixed range, positive or negative, higher for a token the model rates as a more likely next token.

The name comes from statistics, where the logit of a probability pp is its log-odds, log⁡p1−p\log \frac{p}{1-p}, the inverse of the sigmoid function. Neural networks use the word for any raw scores that sigmoid or softmax turns into probabilities, although only a sigmoid's input is exactly a log-odds.

For a sequence of TT tokens, the output layer returns a [T, V] matrix of logits, one row sts_t per position.

Weight tying

GPT-2 has no separate WW. It reuses the transpose of the token embedding matrix EE, [V, d], the matrix in which the embedding layer looks up input ids:

st=htE⊤s_t = h_t E^\top

The output layer adds no parameters of its own, which saves V × d ≈ 38.6M parameters, about 31% of GPT-2 (124M). Weight tying is not universal: many recent large models keep a separate output matrix, while smaller models, where the embedding matrix is a larger share of the parameters, often still tie.

In training, EE gets a gradient from each of its two uses. Through the output layer, all V × d weights get one, because every logit enters the softmax. Through the embedding layer, the input tokens' weights get a second one, added to the first. The optimizer then updates all V × d weights, each with its total gradient.

Training

Softmax turns a row of logits into V probabilities pp, one per vocabulary token, that sum to 1. The loss of one prediction is the cross-entropy. It compares pp with yy, the one-hot encoding of the target: 1 at the token that actually comes next, 0 at every other. Only the target's term is nonzero, so the loss is the negative log of its probability:

loss=−∑k=1Vyklog⁡pk=−log⁡ptarget\text{loss} = -\sum_{k=1}^{V} y_k \log p_k = -\log p_{\text{target}}

Row tt predicts the token at position t+1t + 1, so the targets are the inputs shifted by one position, as in the data pipeline. The causal mask keeps every row from seeing its target, so one forward pass yields TT valid predictions, and an LLM is usually trained on the mean loss over every prediction in the batch.

Inference

Generation uses only the last token's row, the one that predicts a new token. It is autoregressive: the chosen token is appended to the input and the model runs again, one token per step. Choosing the token whose logit has the highest value is greedy decoding. That token also has the highest probability after softmax, so greedy decoding can skip softmax. Sampling instead draws the next token from the probabilities, usually after temperature scaling or top-k filtering of the logits.

References

Sebastian Raschka, Build a Large Language Model (From Scratch) (Manning, 2024). Chapter 4 builds the output layer and generates text with it, chapter 5 computes the loss.

Alec Radford et al., Language Models are Unsupervised Multitask Learners (2019). The model the card runs.

Ofir Press and Lior Wolf, Using the Output Embedding to Improve Language Models (2016). One of the papers that proposed weight tying.

Search

Search pages, articles, and resources