Most levels run in your browser and need nothing installed. A few parts run on your own computer: the bosses that train real models with PyTorch, and a few optional runs. Set this up once; it takes about ten minutes.
Each part that runs on your computer links the files it needs, for example boss.py. Make one course folder, and inside it one folder for each level. Download that level’s files into its folder, and run every command for that level in that folder. The table in step 6 lists every file.
Open a terminal in that folder. A terminal is a window where you type commands. On macOS it is the app Terminal (Applications → Utilities). On Windows use PowerShell (from the Start menu). In the terminal,cd followed by the folder’s path moves you into the folder.
You need Python 3.10 or newer. Check what you have:
python3 --versionInstall from python.org, or with Homebrew: brew install python.
Install from python.org and check the box “Add python.exe to PATH”. Then use py where this page says python3.
Use your package manager, for example sudo apt install python3 python3-venv.
A virtual environment is a private folder of packages for this course, so nothing conflicts with the rest of your computer. Make it once, in a folder that holds your level folders (for example llm-by-hand/):
python3 -m venv .venv
# on Windows use: .venv\Scripts\activate
source .venv/bin/activate
pip install numpy torchWhen the environment is active, your prompt starts with (.venv). In a new terminal, run thesource line again before working on the course.
With the environment active, go into the level’s folder and run the file there. For example, for level N4:
cd lstm # the folder that holds boss.py
python boss.pyThe scripts write their own files (a trained model, downloaded digits) into that folder or into a cache in your home folder.
Everything works on a plain CPU. Some scripts use a faster device when they find one: "cuda" on a computer with an NVIDIA graphics card, "mps" on a Mac with an Apple chip (M1 or newer), otherwise"cpu". You can check what PyTorch sees with:
python -c "import torch; print(torch.cuda.is_available(), torch.backends.mps.is_available())"Times on a recent laptop; a CPU without a GPU can be 2–3× slower.
| level | files | time |
|---|---|---|
| 9 From NumPy to PyTorch | first_run.py | about 1 s |
| U5 Debugging a model | broken_train.py, check.py | a few seconds |
| N4 LSTM and GRU | boss.py | a few minutes |
| N5 Seq2seq and the first attention | boss.py | a few minutes |
| N6 Autoencoders and VAEs | demo.py | about 12 s (downloads the digits once) |
| D3 Latents and DiT | boss.py | about 2–3 min (downloads the digits once) |
| 21 Write your own GPT | skeleton.py, skeleton_hints.py, check.py, check_weights.json | your model trains in about 2 min |
py).source .venv/bin/activate line first.pip install --upgrade pip), and check that your Python is 64-bit and not newer than the newest version PyTorch supports. If it still fails, install the CPU-only build: pip install torch --index-url https://download.pytorch.org/whl/cpu. The browser levels don’t need torch at all.pip install -i <mirror URL> numpy torch, with the address of a PyPI mirror in your country.source line again.