The Math Behind the Machine/ Unit 21 · Build a GPT Checks 0/2 Sign in
Unit 21 of 21 · by Prof. SaurabhFree preview

Build a GPT

A GPT does one thing: it guesses the next letter. Then it adds that letter and guesses again.
In this unit you build one yourself, one part at a time, and see how much better each part makes the guess. Then you train it in your browser and watch it go from gibberish to words in a few minutes.

In the full unit
≈ 180 min read + play 17 interactive widgets · 2 in 3D · a GPT that trains in your browser 24 inline checks 🧾 33 proofs, folded away — open "if you want the algebra" when you are ready ✍ 16 solved practice problems

Section 1 is free. Unlock the other 14 sections and the practice arena for ₹299 — or all eight paid units for ₹999.

Unlock Unit 21 · ₹299Pricing

← Unit 20 · From Noise to Pictures: VAEs and Diffusion

a GPT, assembled · drag to orbit
1

What we will build

Imagine this

A chai stall outside a busy station. You say "One cutting chai and…". The owner is already reaching for the rusk.

Nobody gave him a rule. He has simply heard ten thousand orders. So he knows what usually comes next.

In this unit you build a machine that learns this kind of knowing from text.

The question. Can you build a small ChatGPT on your own laptop, today?

Yes. Not as big as ChatGPT — that needs a building full of computers. But the same kind of machine, built by you, part by part. You will train it in your browser and watch it learn to write.

The one game. A GPT plays one game, again and again:

  1. Read the text so far.
  2. Guess the next letter.
  3. Add that letter to the text. Go back to step 1.

That is all. Writing a whole play is just this game, played many times (Unit 19).

It does not guess one letter. It places bets. After "First Citiz" it might put almost all its bet on "e", and a tiny bit on everything else. Play this yourself below, against the machine you are going to build.

Beat the machineYou have 10 coins. Bet them on the next letter. The machine bets too. Whoever is less surprised wins the round.
Try:
  1. Read the line. Which letter comes next?
  2. Tap a letter to put a coin on it. Tap again for more coins. Very sure? Put all 10 on one letter. Not sure? Spread them.
  3. Press Reveal. The ball on the curve shows how surprised you were. Play a few rounds.

The surprise score. Look at the curve you just played on. It turns "how much did you bet on the right letter" into "how surprised were you":

  • All 10 coins on the right letter: share 1. Surprise 0. You knew it.
  • 5 coins: share ½. Surprise 0.69.
  • 1 coin: share 1/10. Surprise 2.30.
  • 0 coins: surprise with no end. Never bet nothing on a letter that can happen.

The rule behind the curve is short. If you bet a share pp on the letter that really came, your surprise is

surprise=−ln⁡p.\text{surprise}=-\ln p.

Check it: −ln⁡1=0-\ln 1=0, −ln⁡0.5≈0.69-\ln 0.5\approx0.69, −ln⁡0.1≈2.30-\ln 0.1\approx2.30. (Unit 14 met this as cross-entropy.)

The loss. Play many rounds and take the average surprise. That average is called the loss. Lower is better. Training a GPT means one thing: make this number small.

The number to beat. Our text (plays by Shakespeare) uses 65 different characters: letters, capitals, space, new line, punctuation. A machine that knows nothing spreads its bet evenly: 1/651/65 on each. Whatever comes, its surprise is

−ln⁡165=ln⁡65≈4.17.-\ln\tfrac{1}{65}=\ln65\approx4.17.

So 4.17 is the score of knowing nothing. Every real training run starts there. Our job is to push it down.

Where we are goingThe machine you will build, already trained. Give it a start. It writes one letter at a time.
Try:
  1. Press ▶ write. Watch names, new lines and real words appear. Nobody taught it any of them.
  2. The bars below the text are its bets for the next letter, one bar per character. The gold bar is the letter it picked.
  3. Type your own start, like JULIET: or My lord, and write again.
  4. Switch to untrained. All bars are equal, and the text is noise.

The plan: a ladder. We build the machine in six steps. After each step we train it the same way, on the same text, for the same time. Then we measure its loss on text it has never seen.

Every new part must push the loss down. If a part does not help, it does not go in.

The loss ladderEach step down is one new part. Lower is better. The dashed line is the know-nothing score, 4.17.
Try:
  1. Press ▶ walk down the ladder.
  2. Watch the number drop and the text below turn from noise into words, names and lines.
  3. Tap any step to stop there and read what that version writes.
Why does this work?

To bet well, you must know things.

After "KING RICHARD" comes ":" and a new line — because this is a play. After "th" usually comes "e" — because that is English spelling. After "my good" often comes "lord" — because that is how these people talk.

Nobody tells the machine any of this. But knowing it makes the surprise smaller. So when training pushes the surprise down, the machine is forced to pick up spelling, grammar, names and style.

Trap

A machine could just memorise its training text by heart. On that text its loss would be almost 0 — and it would still be lost on anything new.

So we lock away the last tenth of the text and never train on it. Every score in this unit is measured on that locked part.

The realization

loss=average of (−ln⁡pright letter)know nothing: ln⁡65≈4.17\begin{gathered}\text{loss}=\text{average of }\big(-\ln p_{\text{right letter}}\big)\\ \text{know nothing: }\ln 65\approx4.17\end{gathered}

One game: bet on the next letter. One score: the average surprise, on text the machine has never seen. Building a GPT means adding parts that make this score smaller.

Pause & predict

A different text uses 130 characters instead of 65. What does a know-nothing machine score now?

Pause & predict

A machine scores 0.05 on its training text, but 3.9 on the locked-away text. What happened?

If you want the algebra · 2 proofs, step by step
Prove it · a model that knows nothing scores exactly ln V

Claim. A model that gives each of VV characters the probability 1/V1/V has loss ln⁡V\ln V on any text at all. For our 65 characters that is ln⁡65≈4.1744\ln65\approx4.1744.

1
Whatever the true next character is, its probability is 1/V1/V, so its surprise is −ln⁡1V=ln⁡V.-\ln\tfrac1V=\ln V. −ln⁡(1/x)=ln⁡x-\ln(1/x)=\ln x: dividing becomes a minus sign inside the log.
2
Every one of the nn surprises is the same ln⁡V\ln V, so their average is ln⁡V\ln V. ∎ The text does not matter. That is why step 0 of every run should print ln⁡V\ln V: a fresh model starts with tiny random numbers, which give almost equal probabilities.
Prove it · no model can beat the text's own unpredictability

Claim. Suppose the text really comes from probabilities pp (the true chances of each next character), and the model guesses qq. The model's average surprise is the cross-entropy H(p,q)=−∑cpcln⁡qcH(p,q)=-\sum_c p_c\ln q_c, and it is never below the entropy H(p)=−∑cpcln⁡pcH(p)=-\sum_c p_c\ln p_c. It equals it only when q=pq=p.

1
Subtract: H(p,q)−H(p)=∑cpcln⁡pc−∑cpcln⁡qc=∑cpcln⁡pcqc.\begin{aligned}&H(p,q)-H(p)\\ &=\sum_c p_c\ln p_c\\ &\quad-\sum_c p_c\ln q_c\\ &=\sum_c p_c\ln\frac{p_c}{q_c}.\end{aligned} This difference is the KL divergence of Unit 14.
2
Use ln⁡x≤x−1\ln x\le x-1 (the curve ln⁡x\ln x lies under its tangent line at x=1x=1) with x=qc/pcx=q_c/p_c: −∑cpcln⁡qcpc≥−∑cpc(qcpc−1)=−∑cqc+∑cpc=−1+1=0.\begin{aligned}&-\sum_c p_c\ln\frac{q_c}{p_c}\\ &\quad\ge-\sum_c p_c\Big(\frac{q_c}{p_c}-1\Big)\\ &\quad=-\sum_c q_c+\sum_c p_c\\ &\quad=-1+1=0.\end{aligned} Both lists of probabilities add up to 1.
3
So H(p,q)≥H(p)H(p,q)\ge H(p). Equality needs ln⁡x=x−1\ln x=x-1 at every cc, which happens only at x=1x=1: qc=pcq_c=p_c. ∎ Every language has a floor of surprise that no model can get under — some next characters are truly uncertain. Training pushes the loss down toward that floor, never through it.

The road ahead. Four acts.

  1. The goal (§1–§3): turn text into numbers, and build the simplest machine that can play.
  2. Build the machine (§4–§9): one part at a time.
  3. Teach it (§10–§12): one training step by hand, then train it yourself, live.
  4. Use it and open it up (§13–§15): make it write, look inside, keep it.

In one sentence: A GPT bets on the next letter; its score is its average surprise on text it has never seen; knowing nothing scores 4.17, and every part we add must push that down.

Free preview · Unit 21 of 21

That was section 1. The rest of the unit opens when you unlock it.

14 more sections and the practice arena — 17 widgets, 24 checks and 16 solved problems in the whole unit (this preview had 3 widgets and 2 checks).

Unlock Unit 21

  1. 2

    Text becomes numbers

  2. 3

    The first model: counting pairs

  3. 4

    A vector for every character

  4. 5

    Attention: let every character look back

  5. 6

    Many heads, many questions

  6. 7

    Think after you talk: the feed-forward layer

  7. 8

    Going deep: shortcuts and layer norm

  8. 9

    The blueprint, and how big it is

  9. 10

    One training step, every number visible

  10. 11

    Train it yourself

  11. 12

    Reading the loss curve

  12. 13

    Make it write

  13. 14

    Look inside

  14. 15

    Keep it — and the road from here to ChatGPT

  15. 16

    Practice arena — sixteen problems, solved in full

Unlock this unit for ₹299, or all eight paid units for ₹999 — one-time payment, full refund within 7 days. See pricing.