What we will build
A chai stall outside a busy station. You say "One cutting chai and…". The owner is already reaching for the rusk.
Nobody gave him a rule. He has simply heard ten thousand orders. So he knows what usually comes next.
In this unit you build a machine that learns this kind of knowing from text.
The question. Can you build a small ChatGPT on your own laptop, today?
Yes. Not as big as ChatGPT — that needs a building full of computers. But the same kind of machine, built by you, part by part. You will train it in your browser and watch it learn to write.
The one game. A GPT plays one game, again and again:
- Read the text so far.
- Guess the next letter.
- Add that letter to the text. Go back to step 1.
That is all. Writing a whole play is just this game, played many times (Unit 19).
It does not guess one letter. It places bets. After "First Citiz" it might put almost all its bet on "e", and a tiny bit on everything else. Play this yourself below, against the machine you are going to build.
The surprise score. Look at the curve you just played on. It turns "how much did you bet on the right letter" into "how surprised were you":
- All 10 coins on the right letter: share 1. Surprise 0. You knew it.
- 5 coins: share ½. Surprise 0.69.
- 1 coin: share 1/10. Surprise 2.30.
- 0 coins: surprise with no end. Never bet nothing on a letter that can happen.
The rule behind the curve is short. If you bet a share on the letter that really came, your surprise is
Check it: , , . (Unit 14 met this as cross-entropy.)
The loss. Play many rounds and take the average surprise. That average is called the loss. Lower is better. Training a GPT means one thing: make this number small.
The number to beat. Our text (plays by Shakespeare) uses 65 different characters: letters, capitals, space, new line, punctuation. A machine that knows nothing spreads its bet evenly: on each. Whatever comes, its surprise is
So 4.17 is the score of knowing nothing. Every real training run starts there. Our job is to push it down.
The plan: a ladder. We build the machine in six steps. After each step we train it the same way, on the same text, for the same time. Then we measure its loss on text it has never seen.
Every new part must push the loss down. If a part does not help, it does not go in.
To bet well, you must know things.
After "KING RICHARD" comes ":" and a new line — because this is a play. After "th" usually comes "e" — because that is English spelling. After "my good" often comes "lord" — because that is how these people talk.
Nobody tells the machine any of this. But knowing it makes the surprise smaller. So when training pushes the surprise down, the machine is forced to pick up spelling, grammar, names and style.
A machine could just memorise its training text by heart. On that text its loss would be almost 0 — and it would still be lost on anything new.
So we lock away the last tenth of the text and never train on it. Every score in this unit is measured on that locked part.
One game: bet on the next letter. One score: the average surprise, on text the machine has never seen. Building a GPT means adding parts that make this score smaller.
A different text uses 130 characters instead of 65. What does a know-nothing machine score now?
A machine scores 0.05 on its training text, but 3.9 on the locked-away text. What happened?
If you want the algebra · 2 proofs, step by step
Claim. A model that gives each of characters the probability has loss on any text at all. For our 65 characters that is .
Claim. Suppose the text really comes from probabilities (the true chances of each next character), and the model guesses . The model's average surprise is the cross-entropy , and it is never below the entropy . It equals it only when .
The road ahead. Four acts.
- The goal (§1–§3): turn text into numbers, and build the simplest machine that can play.
- Build the machine (§4–§9): one part at a time.
- Teach it (§10–§12): one training step by hand, then train it yourself, live.
- Use it and open it up (§13–§15): make it write, look inside, keep it.
In one sentence: A GPT bets on the next letter; its score is its average surprise on text it has never seen; knowing nothing scores 4.17, and every part we add must push that down.
That was section 1. The rest of the unit opens when you unlock it.
14 more sections and the practice arena — 17 widgets, 24 checks and 16 solved problems in the whole unit (this preview had 3 widgets and 2 checks).
- 2
Text becomes numbers
- 3
The first model: counting pairs
- 4
A vector for every character
- 5
Attention: let every character look back
- 6
Many heads, many questions
- 7
Think after you talk: the feed-forward layer
- 8
Going deep: shortcuts and layer norm
- 9
The blueprint, and how big it is
- 10
One training step, every number visible
- 11
Train it yourself
- 12
Reading the loss curve
- 13
Make it write
- 14
Look inside
- 15
Keep it — and the road from here to ChatGPT
- 16
Practice arena — sixteen problems, solved in full
Unlock this unit for ₹299, or all eight paid units for ₹999 — one-time payment, full refund within 7 days. See pricing.