Level D3 · Theory · Diffusion · runs in your browser · last part on your computer

Latents and DiT

Why does image generation come back to the Transformer?

Side trip · best after level 16 · Transformer parts

Boss level

Write the full sampler yourself and generate a chosen shape in your browser. Then finish boss.py on your computer: your own noise, guidance and reverse step must make a tiny Transformer draw digits that a separate classifier can read.

So far every point had 2 numbers. A small gray picture of a handwritten digit is 28 × 28 = 784 numbers. Nothing in levels D1 or D2 depended on how many numbers a point has. Adding noise, guessing it, stepping back: all the same formulas, now on 784 numbers at once.

What changes is the noise-guessing network. It has to look at a whole picture and say, for every pixel, how much of it is noise. A stroke that starts in one corner continues somewhere else, so every part of the picture needs to know about the other parts. You have already met the tool built for exactly that: attention, from level 14.

What you need: levels D1–D2, and attention and Transformer blocks (levels 14–16). Section 3 uses the autoencoders of side trip N6; it explains what it needs, but N6 shows it in full.

1. A picture as a sentence

A Transformer reads a sequence of tokens. So we cut the picture into square patches and put them in a row, left to right, then top to bottom. Each patch becomes one token: its pixels in a row.

Patch size 2 means each patch is 2 × 2 pixels. Here is a 4 × 4 picture holding the numbers 0 to 15, cut into patches of size 2:

 0  1 |  2  3          token 0: [ 0,  1,  4,  5]
 4  5 |  6  7          token 1: [ 2,  3,  6,  7]
------+------    →     token 2: [ 8,  9, 12, 13]
 8  9 | 10 11          token 3: [10, 11, 14, 15]
12 13 | 14 15

Four patches, four pixels each, so the token matrix is (4, 4).

Careful: “patch size 2” does not mean “cut into a 2 × 2 grid of pieces”. On this 4 × 4 picture both readings happen to give four patches of four pixels, so the example can’t distinguish them. On a 6 × 6 picture they differ. Try it.

Shape A 6 × 6 picture is cut into patches of size 2 (each patch 2 × 2 pixels). What shape is the token matrix (tokens, numbers per token)?

Real digits are bigger, so the same counting gets more interesting.

🔒 Answer the question above to unlock
Number A 28 × 28 digit is cut into patches of size 7 (each patch 7 × 7 pixels). How many tokens is that?
ChooseOn a 28 × 28 picture, go from patch size 4 to patch size 2. What happens to the number of attention scores in each layer?

Now a real digit. Point at or tap a patch to find its token, or a token to find its patch. Change the patch size and watch the token count.

Cut a picture into patches: each patch becomes one token

Pick a patch size. Point at, tap or arrow-key through the patches to see which token each one becomes.

picture, cut into 16 patches of 7 × 7 pixels
tokens: shape (16, 49), one row per patch
Inside a DiT with patch size 7
picture(28, 28)784 numbers
cut into patches(16, 49)16 tokens of 49 numbers
@ W_in(49, 96) → (?, ?)each patch becomes a token vector
+ position(?, ?)which patch is where
+ step t and label(96) added to every tokenlike D1–D2, but as a vector
Transformer blocks(?, ?) → (?, ?)attention scores per head: ? (answer the question below)
@ W_out(96, 49) → (16, 49)a noise guess for every patch
put patches back(28, 28)the noise guess for the picture
patch size
16 patches → 16 tokens of 49 numbers each. Choose a patch to follow it.

Write the cutting yourself, in two steps. First, reshape can split each axis of the picture into (which block, which pixel inside the block). For the 4 × 4 picture and patch size 2:

x4 = img.reshape(2, 2, 2, 2)    # axes: (block row, row in block, block column, column in block)
x4[0, :, 1, :]                  # block row 0, block column 1 → [[2, 3], [6, 7]], the top-right patch

Write the sizes for any picture and patch size:

CodeSplit each axis of an (H, W) picture into (which block, which pixel in the block), for patch size p.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Now one patch is x4[i, :, j, :], but the two block axes, i and j, sit in positions 0 and 2. To list patches in order they must come first. transpose reorders axes: a.transpose(1, 0) swaps the two axes of a table, and x4.transpose(0, 2, 1, 3) gives the new order axis 0, axis 2, axis 1, axis 3, so the shape becomes (block row, block column, row in block, column in block). After that, a reshape to (number of patches, p · p) flattens each patch.

CodeWrite patchify: an (H, W) picture → (number of patches, p·p), patches left to right, then top to bottom.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Once a picture is a sequence of tokens, any Transformer can read it.

🔒 Answer the question above to unlock

2. The DiT: a Transformer that guesses noise

The table under the lab is the whole model. Patches in, a noise guess for every patch out. The middle is the same pre-norm block you built in level 16: attention, feed-forward, LayerNorm, residuals. It has no causal mask, because it writes no sequence one token at a time: every patch may look at every other patch. Only two things are new: the step t and the label (which digit to draw) are turned into vectors and added to every token.

Shape Patch size 4 on a 28 × 28 digit gives 49 tokens of 16 numbers. Wᵢₙ is (16, 96). What is the shape of tokens @ Wᵢₙ?
Number With 49 tokens, how many attention scores does one head compute in one layer?
I got stuck here Why a Transformer? The small network in level D1 worked fine.

In level D1 each point was 2 numbers, and the network could see both at once. A picture is 784 numbers. A plain network that connects every pixel to every pixel needs a huge first layer and does not know that pixel 30 is right below pixel 2.

Patches keep nearby pixels together. Attention then lets every patch ask every other patch “what is drawn over there?” That is exactly the kind of question a noise guess needs: a faint gray pixel is ink if the stroke continues next to it, and noise if it doesn’t.

Go deeper How real models give t and the label to the network

Our tiny model simply adds one vector, made from t and the label, to every token. That works.

Larger models do something stronger. From t and the label they compute a scale and a shift for each LayerNorm inside every block, so the step and the label can make whole features larger or smaller everywhere in the network. Many models also start those scales at zero, so each block begins as “do nothing” and training slowly makes it active.

The rest is plain Transformer: the same blocks, stacked deeper, with more patches.

Patches work for small digits. Large pictures need one more idea.

🔒 Answer the question above to unlock

3. Diffusing in a smaller space

Big pictures make this expensive. A 512 × 512 color picture has 3 numbers per pixel (red, green, blue), so 512 × 512 × 3 = 786,432 numbers. Cut into patches of size 8, it becomes 64 × 64 = 4,096 tokens, and attention compares every one of them with every other. The fix is to squeeze the picture first.

You built the tool for this in level N6: an autoencoder. It is two networks trained together. The encoder squeezes a picture into a few numbers, the code. The decoder turns the code back into a picture. They are trained so the picture that comes back matches the one that went in.

Then you run all of levels D1 and D2 on the codes instead of the pixels, and decode only once, at the very end. The lab below uses a small autoencoder that squeezes each digit into 16 numbers.

Number The autoencoder squeezes 784 pixels into 16 numbers. How many times fewer numbers does diffusion have to work on?
🔒 Answer the question above to unlock

One detail matters before you add noise to codes. Level D1’s formula keeps the variance at 1 only if the clean data has variance 1. The codes from this plain autoencoder are much bigger: their std is about 4, so some codes reach ±10.

Number The 16 code numbers have a std of about 4. Level D1’s formula needs clean data with variance 1, so std 1. You divide every code number by the same number first. Which number?

So the recipe is: encode, divide by the codes’ std, run D1 and D2 on the scaled codes, multiply back, decode. Level N6 also showed why the kind of autoencoder matters. A plain autoencoder leaves holes between codes; a VAE, trained with a small KL term, keeps the codes smooth and close to size 1. Image generators use VAE-style encoders for that reason, so diffusion can land anywhere in the code space and still decode to a sensible picture.

Go deeper What real latents look like

Our code is a flat list of 16 numbers. The encoders inside large image models keep a small grid instead: for example, a 512 × 512 color picture becomes a 64 × 64 grid with 4 numbers per cell. That grid is cut into patches exactly like the pixels in section 1, and a DiT reads those patches as tokens. So the two ideas of this level work together: squeeze first, then patchify the squeezed grid.

Diffusion directly on pixels also works, as the boss below shows; it is just more expensive for big pictures.

The lab also walks from one digit to another: once by mixing pixels, once by mixing the 16-number codes. Guess first.

ChooseAn autoencoder turns each 28 × 28 digit into a 16-number code, and its decoder turns a code back into a picture. Take a 3 and a 7 and go halfway from one to the other. What does mixing the pixels give, compared with mixing the two codes and decoding the result?
🔒 Answer the question above to unlock

Squeeze a picture into 16 numbers, and back

A trained autoencoder keeps 16 numbers per picture. Then walk from one digit to another in two ways.

Loading the trained autoencoder…

loading…
positive code number negative code number
Try it

Slide the walk from the 3 to the 7 slowly. On the left, two faint digits on top of each other. On the right, one digit that changes shape. The 16 numbers describe what is drawn, not where each pixel is. That is why a diffusion model working on codes can make clean pictures with far less work.

You now have every piece. The boss puts them together.

🔒 Answer the question above to unlock

4. Boss, part 1: write the whole sampler

Here is a dataset of 8 points on a ring. For a dataset this small you can write down the perfect noise-guessing function, eps_exact, with no training at all. Your job is the loop that turns random points into points on the ring (the code prints 8 samples; the tests draw 300): the reverse step from level D2, every step from T down to 1, with no fresh noise on the last one.

CodeWrite the sampling loop. eps_exact is the perfect noise guess for this ring of 8 points. Every sample must land on one of the 8 points, and all 8 must appear.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

5. Boss, part 2: draw digits on your computer

This part runs on your computer. New to running Python locally? Start with the setup page. Download boss.py into a folder of its own, and run every command in that folder.

Open boss.py. Three functions are marked TODO, and each is a step you already did in your browser:

  • add_noise: the one line from level D1,
  • guide: the guidance mix from level D2,
  • step_back: the reverse step from level D2, with no fresh noise on the last step. You just wrote it in NumPy in part 1; in PyTorch the fresh noise is torch.randn_like(x).

Everything else is written for you: the digit pictures (downloaded once, about 10 MB), the tiny DiT from section 2, the training loop, and the check.

python boss.py

The script first tests your three functions on the small numbers from levels D1 and D2, so you see a mistake in seconds. Then it trains for STEPS = 8000 steps with your add_noise (about two minutes on our laptop, longer on a slow one), draws 8 pictures of every digit with your guide and step_back, saves them as my_digits.png (one row per digit, 0 at the top), and gives them to a separate digit classifier. You pass the boss when the classifier reads at least 70% of your 80 pictures as the digit you asked for. The script then prints a line that starts with D3 PASS. A correct solution scores about 85–90% on a laptop; the exact number depends on the run. If the small checks pass but the score stays under 70%, open my_digits.png, then raise STEPS near the top of boss.py and run again. Paste the D3 PASS line here. The code at the end of the line only shows that you pasted it unchanged; this part relies on your honesty.

The level counts as cleared once part 1 (the sampler on this page) and this line both pass.

CodeThe boss, part 2. Paste the PASS line that boss.py printed for your own add_noise, guide and step_back, between the quotes.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

There is no reference solution to download: the boss is the test. If you are stuck, go back to the small checks at the top of boss.py. Each of your three functions is one formula from level D1 or D2.

To finish in minutes, this boss diffuses the 784 pixels directly. Doing the same loop on scaled codes from section 3 and decoding at the end is how large image models handle big pictures. The training hides the label 10% of the time; D2’s model hid 20%. Anything from about 10% to 20% works.

The loss falls from about 0.14 to about 0.064 over the 8000 steps. That is far below the 0.32 lower limit from level D1, because that limit depends on the data. Neighboring pixels of a digit predict each other strongly, so a noisy digit shows much more about its noise than a noisy spiral point does.

Then try

Change 3.0 in the call to guide near the end of main to 0, then to 8, and open my_digits.png after each run. At 0 a row shows the wrong digits, and the score drops. At 8 every digit in a row looks the same, and thicker. That is the trade-off between accuracy and variety from level D2, now on pictures.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Cut a picture into patch tokens with reshape and transpose.
  • Count the tokens and attention scores for a picture size and a patch size.
  • Write the whole sampler: reverse steps from T down to 1, with no fresh noise on the last one.

Keep in mind

  • An (H, W) picture with patch size p → (H/p · W/p, p · p); 30 × 30 with p = 5 → (36, 25)
  • Patch tokens in three steps: x.reshape(H // p, p, W // p, p), then .transpose(0, 2, 1, 3), then .reshape(-1, p * p)
  • Attention scores per head and layer = tokens²: halving p gives 4× the tokens and 16× the scores
  • Diffusing on codes: encode, divide by the codes’ std, run D1 and D2, multiply back, decode

Common mistakes

  • Reading “patch size 2” as a 2 × 2 grid of pieces. It means each patch is 2 × 2 pixels.
  • Adding fresh noise on the last step, or forgetting to set x to the mean there.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars