LLM by Hand

Glossary

211 terms, each in a sentence or two, with the level that teaches it. In every level, the first use of a term has a dotted underline: point at it or tap it to see the same card.

A

activationalso: activations

A value computed inside the network for one input, such as a hidden layer’s output. Unlike a parameter, it is computed again for every input and is not saved with the model. Training keeps it in memory until the backward pass has used it.

Taught in level 5: Multilayer perceptron →
activation functionalso: activation

A simple non-linear (curved) function applied after each weighted sum. Without it, stacked layers would become one linear layer.

Taught in level 4: Activation functions →
Adam

The most common optimizer. It combines momentum with a step size per parameter that adapts to how big that parameter’s gradients have been.

Taught in level 7: Optimization →
AdamW

Adam with weight decay done separately: instead of adding the decay to the loss, it shrinks every weight directly after each step.

Taught in level 8: Generalization →
alignmentalso: alignments

In seq2seq attention, which source position each output step looks at most. Drawn as a table of attention weights.

Taught in level N5: Seq2seq + attention →
attentionalso: self-attention

Each token builds a new vector as a weighted average of all the tokens it may look at, with weights the model computes from their similarity.

Taught in level 14: Self-attention →
autoencoder

A network that squeezes its input into a small latent code and then rebuilds the input from that code.

Taught in level N6: Autoencoders →
autograd

Code that records every operation in the forward pass and then computes all the gradients automatically by walking that record backward.

Taught in level U3: Build your own autograd →

B

backpropagationalso: backprop, backward pass

The way to compute every gradient in a network: start at the loss and go backward layer by layer, multiplying local slopes with the chain rule.

Taught in level 6: Backpropagation →
backpropagation through time

Training an RNN by unrolling it over the sequence and backpropagating through every step. Gradients to early steps tend to shrink or grow very large.

Taught in level N3: Recurrent networks →
batchalso: mini-batch

Several examples processed together in one step, stacked along an extra first axis. A batch of 32 sentences of 10 tokens is a (32, 10, …) tensor.

Taught in level 1: Matrices and shapes →
BatchNormalso: batch normalization, batch norm

Normalizing each feature using the mean and std over the batch. Works well for pictures, but depends on the batch size.

Taught in level N2: Residuals and norms →
beam search

Keeping the few best partial sentences at each step instead of just one, and finally choosing the sequence with the best total score.

Taught in level 18: Generation strategies →
bias

A learned number added after the weighted sum. It moves the point where a neuron starts to give a nonzero output.

Taught in level 3: Neurons and loss →
bidirectional RNNalso: bidirectional

Two RNNs, one reading left to right and one right to left, with their states joined. Good for labeling a whole sentence; not usable for writing text one token at a time.

Taught in level N4: LSTM and GRU →
bigram

A pair of neighboring words. A bigram model predicts the next word from the current word alone, using counts.

Taught in level 11: Language models →
binary cross-entropy

The loss for a yes/no answer: −ln p when the target is 1 and −ln(1 − p) when it is 0, averaged over the examples.

Taught in level 3: Neurons and loss →
bit

Here: the unit of surprise when you use log₂. 1 bit is the surprise of an event with probability 1/2.

Taught in level 10: Probability and sampling →
bottleneck

A place where all information must pass through a small vector. In a plain seq2seq model the whole source sentence must fit into one fixed-size state.

Taught in level N5: Seq2seq + attention →
BPEalso: byte-pair encoding

A tokenizer that starts from characters and repeatedly merges the most frequent neighboring pair into a new token.

Taught in level 12: Tokens and BPE →
broadcasting

The rule that lets arrays of different shapes be added or multiplied by stretching size-1 axes to match. A (4, 1) column plus a (1, 5) row gives a (4, 5) table.

Taught in level 1: Matrices and shapes →
bucket

A group of examples with about the same length, batched together so little padding is needed.

Taught in level U4: Feeding data →

C

capacity

How complicated a pattern a model can fit. More parameters give more capacity; too much capacity for the data leads to overfitting.

Taught in level 8: Generalization →
causal mask

A mask that blocks every position from looking at later positions, so a model that writes left to right cannot see the future.

Taught in level 15: Masks and heads →
cell state

The LSTM’s long-term memory, a vector passed along almost unchanged from step to step unless the gates edit it.

Taught in level N4: LSTM and GRU →
chain rule

To find how the loss changes with an early number, multiply the rates of change along the way from that number to the loss.

Taught in level 6: Backpropagation →
channelalso: channels

One of the stacked 2D maps in a convolution input or output, such as red, green and blue in a color picture, or one filter’s feature map.

Taught in level N1: Convolutions →
code

In an autoencoder, the few numbers the encoder outputs (also called the latent). Not program code.

Taught in level N6: Autoencoders →
collate function

The function that turns a list of examples into one batch array, padding the shorter sequences and returning the padding mask with it.

Taught in level U4: Feeding data →
computation graph

The record of which numbers were combined into which, drawn as nodes and arrows. Backpropagation walks it in reverse.

Taught in level U3: Build your own autograd →
concat

Short for concatenate: placing arrays next to each other along one axis. Concat of the heads’ outputs gives one wider table.

Taught in level 15: Masks and heads →
context vector

In seq2seq attention, the weighted average of the encoder states that the decoder uses at one step.

Taught in level N5: Seq2seq + attention →
context window

The largest number of tokens the model can read at once.

Taught in level 12: Tokens and BPE →
contiguous

Stored in plain row order: reading the array row by row walks through memory one step at a time. `.contiguous()` makes a copy in that order.

Taught in level U1: Tensors in memory →
convolution

Sliding a small filter over a picture and taking a dot product at each position, so the same pattern can be found anywhere.

Taught in level N1: Convolutions →
cosine similarity

The dot product of two vectors divided by both lengths. It depends only on the angle: 1 for the same direction, 0 for a right angle, −1 for opposite directions.

Taught in level 13: Embeddings and similarity →
cross-attention

Attention where the queries come from one sequence and the keys and values from another, such as the decoder reading the encoder’s memory.

Taught in level N7: 2017 Transformer →
cross-entropy

The usual loss for picking one right answer among many. It is −log of the probability the model gave to the correct answer.

Taught in level 10: Probability and sampling →
crossover

The number of users B at which the arithmetic of one decode step takes as long as reading the weights. Below it, reading is the limit; above it, the arithmetic is. It depends on the GPU and the bytes per weight, not on the size of the model.

Taught in level 20: Inference cost →

D

data pipeline

The code that turns raw examples into batches for training: shuffle, split, group into batches and pad to the same length.

Taught in level U4: Feeding data →
DataLoaderalso: Dataset

PyTorch’s two parts of a data pipeline. A Dataset gives example i; a DataLoader groups examples into batches, one batch at a time.

Taught in level U4: Feeding data →
dead ReLU

A ReLU neuron whose input is below 0 for every example. Its output and its slope are always 0, so its weights stop learning.

Taught in level 4: Activation functions →
decision boundary

The line where a classifier switches from one answer to the other. For one neuron it is the straight line where z = 0.

Taught in level 3: Neurons and loss →
decodealso: decoding

Writing the answer one token per step. Each step reads every weight and the KV cache once for one new token per user, so memory bandwidth, not arithmetic, usually limits it.

Taught in level 20: Inference cost →
decoder

A stack of blocks with masked self-attention that writes the output one token at a time, looking only at tokens already written. In the original two-stack design it also reads the encoder’s memory.

Taught in level 16: Transformer parts →
decoder blockalso: decoder blocks

Masked self-attention, then the FFN, each with its own residual and LayerNorm. A decoder-only model is a stack of these blocks.

Taught in level 16: Transformer parts →
decoder-only

A Transformer with only the decoder stack and masked self-attention. Today’s large language models are built this way.

Taught in level 17: The whole Transformer →
depth-first walk

Visiting a graph by following one path to its end before trying the next. Autograd uses it to build the topological order.

Taught in level U3: Build your own autograd →
diffusionalso: diffusion model

A way to generate data by learning to undo noise: add noise to examples step by step, then train a network to remove it.

Taught in level D1: Adding and removing noise →
DiT

A diffusion Transformer: it cuts a noisy picture into patch tokens and uses Transformer layers to guess the noise.

Taught in level D3: Latents and DiT →
dot product

Multiply two vectors number by number and add up. It is large when the vectors point the same way, near 0 when they are unrelated.

Taught in level 13: Embeddings and similarity →
dropoutalso: inverted dropout

During training, setting some units to 0 at random on each step, so the network can’t rely on any single one. The kept units are divided by 1 − p (inverted dropout), so nothing changes at test time.

Taught in level 8: Generalization →

E

early stopping

Stopping training when the loss on held-out data stops improving, even if the training loss is still going down.

Taught in level 8: Generalization →
embedding

The list of numbers that stands for a token. The embedding table has one learned row per vocabulary entry.

Taught in level 13: Embeddings and similarity →
encoder

The part of the original two-stack Transformer that reads the whole source sentence at once and produces its vectors (the memory).

Taught in level N7: 2017 Transformer →
end tokenalso: <eos>

A special token, written <eos> (“end of sequence”), that ends a sequence. The model writes it when its answer is complete, so generation knows when to stop. Used in level N7 and level 21.

Taught in level 21: Write your own GPT →
entropy

How uncertain a probability distribution is, measured as the average surprise of its outcomes.

Taught in level 10: Probability and sampling →
epochalso: epochs

One pass of training over the whole training set.

Taught in level 7: Optimization →
exploding gradientalso: exploding gradients

When gradients get multiplied by numbers bigger than 1 at every step, so they grow very large and training fails.

Taught in level N3: Recurrent networks →
exponent

In a float, the bits that say which power of 2 the number is multiplied by. More exponent bits mean a wider range, from tiny to huge.

Taught in level U2: Numbers in a computer →
exponent bias

A fixed number added to the power of 2 before it is stored in a float, so that negative powers need no sign. Not the bias of a neuron.

Taught in level U2: Numbers in a computer →

F

fan-in

The number of inputs that feed one neuron (n). Good starting weights are scaled by it, for example std √(2/n) for ReLU.

Taught in level 7: Optimization →
feature map

The output of one convolution filter over a whole picture: a grid that is bright wherever the filter’s pattern appears.

Taught in level N1: Convolutions →
feed-forward networkalso: FFN, position-wise feed-forward network

The small MLP applied to every token separately, after attention, inside each Transformer layer. Because it works on each position separately, it is also called position-wise.

Taught in level 16: Transformer parts →
filteralso: kernel

The small grid of learned weights a convolution slides over the input, such as a 3×3 edge detector.

Taught in level N1: Convolutions →
FLOPs

Floating-point operations: the number of multiplications and additions a computation needs. A way to count how much arithmetic a model needs.

Taught in level 20: Inference cost →
forget gate

In an LSTM, the gate f (between 0 and 1) that decides how much of the old cell state c to keep at each step.

Taught in level N4: LSTM and GRU →
forward pass

Running the inputs through the network from start to end to get the output and the loss.

Taught in level 6: Backpropagation →
fused kernelalso: fused kernels, fused attention kernel, fused attention kernels

One GPU program that does several steps, such as the scores, the softmax and the weighted sum, without saving the values between them in GPU memory.

Taught in level U6: Attention, backward →

G

gate

A sigmoid between 0 and 1 that multiplies a signal, deciding how much of it to keep (1 keeps all of it, 0 keeps none).

Taught in level N4: LSTM and GRU →
Gaussianalso: Gaussian noise, normal distribution

The bell-shaped distribution. Gaussian noise means random numbers that are usually near the mean and rarely far from it.

Taught in level 10: Probability and sampling →
GELU

A smooth version of ReLU that lets small negative values through a little. Most modern language models use it or a similar function.

Taught in level 4: Activation functions →
generalization

How well a model works on new data it has never seen, not only on the examples it trained on.

Taught in level 8: Generalization →
geometric mean

The n-th root of the product of n numbers. The geometric mean of 2 and 8 is √(2 · 8) = 4.

Taught in level 10: Probability and sampling →
GPT

A decoder-only Transformer trained to predict the next token. Level 21 builds a tiny one that learns two-digit addition.

Taught in level 21: Write your own GPT →
GQAalso: grouped-query attention

Attention where several query heads share one key head and one value head, which saves GPU memory when generating.

Taught in level 20: Inference cost →
gradient

The list of slopes of the loss, one per parameter. It points in the direction where the loss grows fastest, so training steps the other way.

Taught in level 2: Gradient descent →
gradient check

Testing a computed gradient by changing one input a tiny bit and measuring how much the loss moves. The two numbers should match.

Taught in level 6: Backpropagation →
gradient descent

Repeatedly changing every parameter a small step against its gradient, so the loss goes down.

Taught in level 2: Gradient descent →
greedy decodingalso: greedy

Generating text by always taking the single most likely next token.

Taught in level 17: The whole Transformer →
GRU

A simpler version of the LSTM with two gates (the update gate and the reset gate) and no separate cell state.

Taught in level N4: LSTM and GRU →
guidancealso: classifier-free guidance

Pushing generation toward what was asked for by comparing the network’s noise guess with and without the request and moving further in that direction.

Taught in level D2: Sampling and guidance →

H

held-out setalso: held-out data, validation set, test set

Examples held out and never trained on, used only to measure how well the model works on data it has not seen.

Taught in level 8: Generalization →
hidden layer

A layer between the input and the output. Its numbers are not given by the data; the network learns what they should mean.

Taught in level 5: Multilayer perceptron →
hidden state

The vector an RNN carries from one step to the next. Everything it remembers about the sequence so far has to fit in it.

Taught in level N3: Recurrent networks →
hidden unit

One neuron in a hidden layer, and the one number it outputs. A hidden layer with 8 hidden units outputs 8 numbers.

Taught in level 5: Multilayer perceptron →

I

idalso: token id

The number that stands for one token: 0 up to the vocabulary size minus one.

Taught in level 12: Tokens and BPE →
in place

Changing an array’s own numbers instead of making a new array. `w[i] += h` changes w in place.

Taught in level U5: Debugging →
inference

Using a trained model to produce outputs. The weights stay fixed; only the activations are computed again for each new input.

Taught in level 14: Self-attention →
initialization

How the weights are set before training starts. Too large and the layer outputs grow very large through the layers; too small and they shrink toward 0.

Taught in level 7: Optimization →
input gate

In an LSTM, the gate i (between 0 and 1) that decides how much of the new candidate g to add to the cell state.

Taught in level N4: LSTM and GRU →

J

Jacobian

The table of every slope between the inputs and the outputs of one operation: row i, column j says how fast output i moves when input j moves. For softmax it is diag(p) − pᵀp.

Taught in level U6: Attention, backward →

K

keyalso: keys

In attention, the vector that other words compare their query against. A query matches keys that point the same way.

Taught in level 14: Self-attention →
KL divergencealso: KL term

A number that measures how different one probability distribution is from another. A VAE uses it to keep its codes close to a standard Gaussian.

Taught in level N6: Autoencoders →
KV cachealso: KV-cache

Saving the keys and values of earlier tokens during generation, so each new token only computes its own query, key and value.

Taught in level 18: Generation strategies →

L

label smoothing

Training with a target that gives the right token slightly less than 100% and spreads the rest over other tokens, so the model doesn’t become too certain.

Taught in level N7: 2017 Transformer →
language model

A model that gives a probability to every possible next token, given the tokens so far. Writing text means picking a token and repeating.

Taught in level 11: Language models →
latentalso: latent space

A short code that stands for a much larger input, such as 16 numbers for a 784-pixel picture. Diffusing in latent space is much cheaper.

Taught in level N6: Autoencoders →
layer

Several neurons side by side that all read the same inputs. A layer turns a (batch, n_in) array into a (batch, n_out) array.

Taught in level 4: Activation functions →
LayerNorm

Rescaling each token’s vector to mean 0 and std 1, then applying a learned gain and shift (γ and β). It keeps numbers in a stable range.

Taught in level 16: Transformer parts →
learning rate

How big each gradient-descent step is. Too small and training is very slow; too big and the loss goes up and down from step to step, or grows very large.

Taught in level 2: Gradient descent →
learning-rate decayalso: decay

Lowering the learning rate as training goes on, so the noisy mini-batch steps end close to the lowest point.

Taught in level 7: Optimization →
learning-rate schedule

A rule that changes the learning rate during training, usually rising for a short warmup and then falling slowly toward 0.

Taught in level 7: Optimization →
local gradient

How one operation’s output moves when its input moves, found for that operation alone. The chain rule multiplies the local gradients along the path.

Taught in level 6: Backpropagation →
log-sum-exp

ln of the sum of e to the power of each score. Subtracting the largest score first keeps it from overflowing; cross-entropy is log-sum-exp minus the target’s score.

Taught in level U2: Numbers in a computer →
logits

The raw scores the model gives each vocabulary token before softmax. Larger means more likely.

Taught in level 17: The whole Transformer →
loss

One number that says how wrong the model is on the data. Training means making this number smaller.

Taught in level 2: Gradient descent →
loss scaling

Multiplying the loss by a big number (such as 1024) before the backward pass, so tiny 16-bit gradients do not become 0, then dividing the gradients by the same number.

Taught in level U2: Numbers in a computer →
LSTM

A recurrent network with gates and a separate cell state, which lets it keep information over many more steps than a plain RNN.

Taught in level N4: LSTM and GRU →

M

mantissa

The bits of a float that hold its digits: the part after “1.” in 1.5 × 2². More mantissa bits mean finer steps between neighboring numbers.

Taught in level U2: Numbers in a computer →
mask

A table that blocks some attention scores before softmax, for example so a token can’t look at tokens that come after it, or at padding.

Taught in level 15: Masks and heads →
master copy

In mixed-precision training, the float32 copy of the weights. Each update is added there, so small changes are not rounded away; the 16-bit weights are made from it.

Taught in level U2: Numbers in a computer →
matrixalso: matrices

A table of numbers with rows and columns. Almost every step inside a neural network is a matrix multiplication.

Taught in level 1: Matrices and shapes →
matrix multiplicationalso: matmul

Each cell of the result is a row of the left matrix times a column of the right one, multiplied number by number and added up. (n, k) @ (k, m) gives (n, m).

Taught in level 1: Matrices and shapes →
mean squared erroralso: MSE

A loss: subtract the target from the prediction, square each difference, and take the average.

Taught in level N6: Autoencoders →
memory

In the encoder–decoder Transformer, the encoder’s final output: one vector per source token. Every decoder layer reads it through cross-attention. (Not an RNN’s hidden state.)

Taught in level N7: 2017 Transformer →
memory bandwidthalso: bandwidth

How many bytes per second a GPU can read from its memory. When a model writes one token at a time, every weight is read once per token, so the bandwidth often limits the speed more than the arithmetic does.

Taught in level 20: Inference cost →
mixed precision

Training with most multiplications in 16-bit numbers for speed, while the weights, sums and the loss stay in 32-bit numbers.

Taught in level U2: Numbers in a computer →
MLPalso: multilayer perceptron

A network made of layers of neurons, each layer a matrix multiplication followed by an activation function.

Taught in level 5: Multilayer perceptron →
momentum

An optimizer trick that keeps a running average of past gradients, so steps get faster where the slope keeps the same direction and stop going back and forth.

Taught in level 7: Optimization →
multi-head attentionalso: head, heads, attention head

Running several smaller attentions side by side, each with its own Q, K and V, then joining their outputs. Each head can track a different relation.

Taught in level 15: Masks and heads →

N

NaN

“Not a number”: the result of a calculation with no answer, like inf − inf or 0 × inf. Once a NaN appears, everything computed from it is NaN too.

Taught in level U5: Debugging →
nat

The unit of surprise when you use the natural log ln. 1 nat is the surprise of an event with probability 1/e.

Taught in level 10: Probability and sampling →
neuron

The basic unit of a network. It takes a weighted sum of its inputs, adds a bias, and passes the result through an activation function.

Taught in level 3: Neurons and loss →
noise schedulealso: alpha-bar

The plan for how much noise each step adds. Alpha-bar says how much of the original signal is left after t steps.

Taught in level D1: Adding and removing noise →
normalizationalso: normalize, normalizes, normalized, normalizing

Rescaling numbers so they stay in a stable range: subtract the mean, divide by the standard deviation. BatchNorm and LayerNorm differ only in which numbers they average over.

Taught in level N2: Residuals and norms →
numeric gradient

A gradient found with no calculus: move one weight by a tiny h, run the forward pass again, and divide the change in the loss by h.

Taught in level 6: Backpropagation →

O

one-hot

A vector that is all zeros except a single 1 at a token’s id. Multiplying it by the embedding table gives exactly that token’s row.

Taught in level 13: Embeddings and similarity →
optimizer

The rule that turns gradients into parameter updates. Plain gradient descent, momentum and Adam are all optimizers.

Taught in level 7: Optimization →
output gate

In an LSTM, the gate o (between 0 and 1) that decides how much of the cell state to show as the output h.

Taught in level N4: LSTM and GRU →
overfitting

When a model memorizes its training examples instead of learning the pattern, so it does well on them and badly on new data. Overfitting one small batch on purpose is a different thing: a quick test that the training code works (level U5).

Taught in level 8: Generalization →
overflow

When a result is bigger than the largest number a format can hold, so it becomes inf. In float16, e¹² already overflows.

Taught in level U2: Numbers in a computer →

P

padding

Extra placeholder tokens (often written <PAD>) added to the end of short sequences so every sequence in a batch has the same length. Masks and the loss ignore them.

Taught in level 15: Masks and heads →
padding (convolution)also: zero padding

Extra border cells (usually zeros) added around an input so the filter can sit on the edges and the output keeps its size.

Taught in level N1: Convolutions →
padding mask

A mask that blocks every PAD token (the extra tokens that make short sequences as long as the longest one), so no word looks at padding.

Taught in level 15: Masks and heads →
parameter

Any number the model learns during training, such as a weight or a bias. Model size is usually counted in parameters.

Taught in level 5: Multilayer perceptron →
patch

A small square cut from a picture. A Transformer for images treats each patch as one token.

Taught in level D3: Latents and DiT →
perplexity

e to the power of the average cross-entropy. A perplexity of 4 means the model is as unsure as if it were choosing among 4 equally likely words.

Taught in level 10: Probability and sampling →
plateaualso: plateaus

A part of training where the loss or accuracy stays flat for a while before it starts to improve again.

Taught in level 21: Write your own GPT →
pooling

Shrinking a feature map by keeping one number (often the largest) from each small block.

Taught in level N1: Convolutions →
positional encoding

Numbers added to each token’s embedding that say where it is in the sequence. Attention alone does not know the word order.

Taught in level 16: Transformer parts →
post-norm

The 2017 order: add the sublayer’s output to x, then normalize, LayerNorm(x + f(x)). Deep post-norm stacks need a long warmup to train.

Taught in level 19: Modern blocks →
pre-norm

Putting the normalization before each sublayer instead of after it, so nothing rescales the residual path and deep stacks train reliably.

Taught in level 16: Transformer parts →
precision

How exact a float format is: how many numbers it has between one power of 2 and the next. More mantissa bits give more precision.

Taught in level U2: Numbers in a computer →
prefill

Reading the whole prompt in one forward pass before the answer starts. The weights are read once for all prompt tokens. So for a long prompt, arithmetic limits it, not memory: on the example GPU of level 20, longer than about 100 tokens. It also fills the KV cache.

Taught in level 20: Inference cost →
prompt

The tokens you give the model at the start. The model continues them one token at a time, and the tokens it writes are the answer.

Taught in level 17: The whole Transformer →

Q

queryalso: queries

In attention, the vector a token uses to ask “what am I looking for?” Scores are queries dotted with keys.

Taught in level 14: Self-attention →

R

random generator

An object that makes random numbers, starting from a number called the seed. The same seed gives the same numbers.

Taught in level U4: Feeding data →
range

How big and how small a number a float format can hold. More exponent bits give a wider range.

Taught in level U2: Numbers in a computer →
receptive field

The patch of the original input that one output number can see. It grows as convolution layers are stacked.

Taught in level N1: Convolutions →
recomputation

Keeping only each layer’s input during the forward pass, and running that layer’s forward pass again during backward to get the other values back. It saves GPU memory at the cost of about one extra forward pass.

Taught in level U6: Attention, backward →
regularization

Any method that stops a model from memorizing the noise in its training data, such as weight decay or dropout.

Taught in level 8: Generalization →
ReLU

The activation max(0, x). It passes positive numbers through and turns negative numbers into 0.

Taught in level 4: Activation functions →
reparameterization trick

Writing a random code as z = μ + σ·ε with ε drawn from a standard Gaussian, so gradients can flow through μ and σ.

Taught in level N6: Autoencoders →
repetition penalty

Lowering the scores of tokens that were already used, so generation doesn’t repeat the same words again and again.

Taught in level 18: Generation strategies →
reset gate

In a GRU, the gate r that decides how much of the old hidden state to use when making the new candidate. With r = 0 the candidate uses only the input.

Taught in level N4: LSTM and GRU →
residual connectionalso: residual, residuals

Adding a layer’s input back to its output, so each layer only has to learn a change, and the values and the gradients pass easily through deep stacks.

Taught in level 16: Transformer parts →
residual pathalso: residual stream

The plain x that goes directly from the input to the output of the stack. Each layer beside it (in a Transformer, each sublayer) reads x and adds its output to x: x + f(x). Gradients go back along this path without being made smaller.

Taught in level 16: Transformer parts →
RMSNorm

A simpler LayerNorm that only divides by the root-mean-square of the vector and skips subtracting the mean.

Taught in level 19: Modern blocks →
RNNalso: recurrent network, recurrent neural network

A network that reads a sequence one step at a time, carrying a hidden state forward and using the same weights at every step.

Taught in level N3: Recurrent networks →
RoPEalso: rotary position embedding

Encoding position by rotating pairs of numbers in the queries and keys, so attention scores depend on how far apart two tokens are.

Taught in level 19: Modern blocks →
running average

An average that is updated at every step: keep most of the old value and add a little of the newest value. Adam keeps two of them.

Taught in level 7: Optimization →
running mean

In BatchNorm, an average of the batch means that is updated after each batch (keep 0.9 of the old value, add 0.1 of the new one). It is used instead of the batch mean when the model runs on new data.

Taught in level N2: Residuals and norms →

S

sampling

Picking at random according to probabilities. A language model samples the next token from its softmax; a diffusion model samples a new picture by starting from noise and running the denoising steps.

Taught in level 10: Probability and sampling →
saturation

When a sigmoid or tanh input is far from 0, the curve is almost flat there, so its slope is nearly 0 and almost no gradient passes back.

Taught in level 4: Activation functions →
saved activationsalso: saved activation

Values from the forward pass that training keeps in GPU memory because a backward rule needs them later, such as the attention weights A. They grow with the batch size and the sequence length.

Taught in level U6: Attention, backward →
seedalso: random seed

The number a random generator starts from. The same seed gives the same “random” numbers every time, so a run can be repeated exactly.

Taught in level U4: Feeding data →
self-attention

Attention where the queries, keys and values all come from the same sequence, so every token looks at the tokens of its own sentence.

Taught in level 14: Self-attention →
seq2seqalso: sequence-to-sequence

A model that reads one sequence and writes another, such as digits in and words out: an encoder reads, a decoder writes.

Taught in level N5: Seq2seq + attention →
shape

The size of an array along each axis, written like (3, 4) for 3 rows and 4 columns. Most bugs in model code are shapes that don’t line up.

Taught in level 1: Matrices and shapes →
sigmoid

An activation that squeezes any number into the range 0 to 1, along an S-shaped curve. σ(0) = 0.5.

Taught in level 3: Neurons and loss →
SiLU

The activation z · sigmoid(z). It looks like ReLU but is smooth and has a small dip below 0.

Taught in level 19: Modern blocks →
slope

How fast a function’s output changes when its input grows a little: the output change divided by the input change. Training is driven by slopes.

Taught in level 4: Activation functions →
smoothing

Adding a small number (often 1) to every count before dividing, so a pair that was never seen still gets a small probability instead of 0.

Taught in level 11: Language models →
softmax

Turns a list of scores into probabilities that are all positive and add up to 1. A bigger score gets a bigger probability.

Taught in level 10: Probability and sampling →
source sentencealso: source sequence

In translation, the sentence the model is given. The encoder reads it. In level N7 it is the digits, such as 427.

Taught in level N7: 2017 Transformer →
standard deviationalso: std

How far numbers typically are from their mean: the square root of the variance. Written std or σ.

Taught in level 10: Probability and sampling →
start tokenalso: <bos>

A special token, written <bos> (“beginning of sequence”), that starts a sequence. It gives the model a first input before any word is written. Used in level N7 and level 21.

Taught in level N7: 2017 Transformer →
steepness

In this level’s lab, the number k in sigmoid(k · z). A bigger k makes the S-curve go from 0 to 1 faster.

Taught in level 3: Neurons and loss →
stride

How many pixels a convolution filter moves between positions. A stride of 2 halves the output size.

Taught in level N1: Convolutions →
stride (memory)also: stride

How many places to move in memory for one step along an axis. A (3, 4) array in row order has strides (4, 1): one row down is 4 places, one column right is 1.

Taught in level U1: Tensors in memory →
sublayeralso: sublayers

One of the parts inside a Transformer block that gets its own residual: the attention or the FFN. A decoder block has two sublayers. A decoder layer of the 2017 design has three.

Taught in level 16: Transformer parts →
surprise

−ln p for the word that actually came. A probability of 1 gives surprise 0; a small probability gives a large surprise. Cross-entropy is the average surprise.

Taught in level 10: Probability and sampling →
SwiGLU

A feed-forward network with a gate: one matrix product of x (the gate, after SiLU) decides how much of another matrix product (the value) passes.

Taught in level 19: Modern blocks →

T

tanh

An S-shaped activation like sigmoid, but centered on 0, with outputs between −1 and 1.

Taught in level 4: Activation functions →
target sentencealso: target sequence

In translation, the sentence the model must write. The decoder writes it. This is not the same as the training target y: tgt_in and tgt_out are both made from the target sentence.

Taught in level N7: 2017 Transformer →
teacher forcing

Training a decoder by giving it the correct previous tokens as input, instead of its own guesses, so every position can be trained at once.

Taught in level 17: The whole Transformer →
temperature

A number the scores are divided by before softmax. Low temperature makes the top choice dominate; high temperature makes the probabilities more even.

Taught in level 10: Probability and sampling →
tensor

An array with any number of axes. A number, a list, a table and a stack of tables are all tensors; PyTorch’s main data type is the tensor.

Taught in level 9: From NumPy to PyTorch →
token

One piece of text the model reads or writes: a character, a word, or a word fragment, depending on the tokenizer. In image models, one patch of a picture.

Taught in level 12: Tokens and BPE →
tokenizer

The rule that cuts text into tokens and maps each token to an id number.

Taught in level 12: Tokens and BPE →
top-k

A sampling rule that keeps only the k most likely tokens and picks among them.

Taught in level 18: Generation strategies →
top-palso: nucleus sampling

A sampling rule that keeps the most likely tokens until their probabilities add up to p, and picks among those.

Taught in level 18: Generation strategies →
topological order

A list of the nodes of a computation graph where each node comes after all of its inputs. Walking it backwards runs the backward pass in a valid order.

Taught in level U3: Build your own autograd →
translation

Turning a sentence in one language into a sentence in another language. The first Transformer was built for this task.

Taught in level N7: 2017 Transformer →
transpose

Flipping a matrix so its rows become columns. A (3, 2) matrix becomes (2, 3); in code it is written `X.T`.

Taught in level 1: Matrices and shapes →

U

underflow

When a result is closer to 0 than the smallest number a format can hold, so it becomes exactly 0.

Taught in level U2: Numbers in a computer →
unrollingalso: unroll, unrolled

Drawing an RNN as one copy per time step, all copies sharing the same weights, so you can see the whole computation as one long network.

Taught in level N3: Recurrent networks →
update gate

In a GRU, the gate z that decides how much of the new candidate to take; the rest, 1 − z, is the old hidden state kept as it was.

Taught in level N4: LSTM and GRU →

V

VAEalso: variational autoencoder

An autoencoder whose encoder outputs a mean and a std for each code number, trained so that random codes decode into plausible examples.

Taught in level N6: Autoencoders →
valuealso: values

In attention, what a token gives to the tokens that look at it. The output is a weighted average of values.

Taught in level 14: Self-attention →
vanishing gradientalso: vanishing gradients

Gradients that shrink layer after layer (or step after step) until the early parts of a network barely learn.

Taught in level 4: Activation functions →
variance

The average squared distance of numbers from their mean. It measures how spread out they are; its square root is the standard deviation.

Taught in level 10: Probability and sampling →
vectoralso: vectors

A list of numbers in a fixed order, such as [0.5, −1, 2]. A word’s vector is the row of numbers that stands for it.

Taught in level 13: Embeddings and similarity →
view

An array that reads another array’s memory with its own shape and strides. Nothing is copied, so writing into a view changes the original.

Taught in level U1: Tensors in memory →
vocabulary

The full list of tokens a model knows. Each one has an id, and the model’s output has one score per vocabulary entry.

Taught in level 12: Tokens and BPE →

W

warmup

The first part of a learning-rate schedule: start with a tiny learning rate and raise it over the first few hundred steps.

Taught in level 7: Optimization →
weight

A number the network learns that says how strongly one input counts. Weights and biases together are the network’s parameters.

Taught in level 3: Neurons and loss →
weight decay

At each step, every weight is made a little smaller, closer to 0. It discourages extreme weights and helps against overfitting.

Taught in level 8: Generalization →
weight sharing

Using the same weights at many positions. A convolution uses the same 9 filter weights at every position of the picture.

Taught in level N1: Convolutions →
word vectoralso: word vectors, word2vec

An embedding learned from which words appear near each other, so words used in similar places get similar vectors.

Taught in level 13: Embeddings and similarity →