Ask something to predict what comes next and it has no choice but to build whatever the next moment depends on. Jeffrey Elman showed it with words in 1990. Claude Shannon had already shown that the same loop, read the other way, is compression.
Take a photograph of a ball in mid-air and ask where it will be a moment later. You cannot say. Up, down, sideways, or barely moving at all, the picture looks the same. Everything the answer depends on is missing from a single frame.
Now take two photographs a fraction of a second apart. This time the answer arrives before you have finished asking. Both pictures were equally sharp, so detail is not what changed.
Two frames hold something one frame cannot: a direction and a speed. Neither is drawn on either picture. They live in the relationship between the two, and you pulled them out without being told to.
What you can work out
Where it is. That is all a single picture can tell you.
What is still open
Every direction, and every speed. The picture rules out nothing.
Ask something to predict what comes next, and it has no choice but to build whatever the next moment depends on. It was never asked for those things, and it cannot be right without them. I call them stowaways: the quantities a predictor is forced to carry that nobody put on the manifest. Direction and speed are the first two.
People have been finding larger stowaways ever since, with better instruments each time. Claude Shannon went first, in 1948, with A Mathematical Theory of Communication. He showed that how well you can predict a message sets the price of sending it.
Jeffrey Elman followed in 1990 with Finding Structure in Time. He found grammar inside a network asked only to guess the next word. Then in 2022 Kenneth Li and colleagues wrote Emergent World Representations, and found a whole game board inside a network asked only to guess moves.
| Year and who | Asked to | Instrument | Had to carry |
|---|---|---|---|
| Opening: the ball | say where the ball will be | two photographs | |
| 1948, Claude Shannon | send a message cheaply | a pencil and probabilities (the maths of information) | |
| 1990, Jeffrey Elman | guess the next word | look at the network's internal states and group them | |
| 2022, Kenneth Li and colleagues | guess the next legal Othello move | a probe trained to read one fact out of the activations |
The manifest lists what each one was asked for. Nothing else is written down.
Rows open
0 of 4
"Prediction is learning" gets said a lot, usually as a slogan, and rarely with any account of what the predictor learned or how anybody checked.
Guess, check, adjust
The loop is as small as it sounds. Look at what you have and say what you think is coming. Wait, then compare what arrived to what you said. Change yourself a little in the direction that would have been less wrong, and go again.
None of the steps is clever. The power is in what the loop does not need: nobody has to label anything, because the target is the next piece of the recording. Let's watch it run.
Recently wrong by
0.426
What it has learned
0.52, 0.17
How it is doing
Guessing. It has seen almost nothing yet.
Nothing in that figure was told the rule. It was told, two hundred and twenty times, by how much it had just missed. Each time it nudged its own two numbers (the weights: the adjustable numbers inside a model, and the only thing that learning ever changes) a little way towards being right, and that was enough. Splendid.
For most of the field's history a model learned from examples that somebody had labelled. A million images meant a person writing down what was in each one. The labels were the costly part, and the size of a dataset was the size of somebody's budget. Prediction skips the bill, because any recording of anything is already a training set.
The argument was still open in 2013, when Yoshua Bengio, Aaron Courville and Pascal Vincent wrote Representation Learning: A Review and New Perspectives. It is a survey, and it reads as a case still being made. That the data was free settled it, and that convenience, more than any single result, is why prediction took over the field.
Labelled by a person
Labelled by the recording
Every example cost somebody a label. The dataset is the size of the budget.
Examples
1k
Labels a person wrote
1k
Who paid
somebody's budget
What guessing forces you to build
Elman was a cognitive scientist at the University of California, San Diego, who came to networks from the study of language. In 1990 he trained a small network on one job: read a word and predict the next word.
Then he looked inside. The internal states had sorted themselves into groups. The groups were nouns and verbs, animate and inanimate, things that can eat and things that can be eaten.
Nobody had given the network a single category. It reached them because a word's category is what its next word depends on. The title, Finding Structure in Time, is literal: the structure was found rather than installed.
Twenty-four words, placed at random. Nothing to read yet.
Hover or tap a dot, or focus the figure and use the arrow keys.
Groups you can see
0
Word
none yet
For a long time that was where it stopped. The network was small and the grammar was a toy, and a toy is not a claim about the world. The finding waited thirty-two years for instruments that could make it one.
In 2022 Li and colleagues, interpretability researchers (people who open networks up rather than train them), took a network since nicknamed Othello-GPT. It was trained only to predict legal moves in Othello (a board game on an eight-by-eight grid. You flip your opponent's discs by trapping them between two of yours). It was never shown a board and never told the rules.
They probed its activations (train a small separate classifier to read one specific fact out of a network's internal numbers). The state of the board was in there, which on its own is a curiosity.
The intervention is what made it a finding. Change that internal board, and the moves the network goes on to make change with it. So the board is there because the network is using it. The two steps are worth keeping apart, because only the second one settles anything.
It plays from a board
The probe works it out
Two stories about one network. Run a test and see which of them it separates.
Test
nothing run yet
Stories still standing
2
What ruled one out
nothing yet
Neel Nanda and colleagues studied the same network again in 2023, in Othello-GPT has a linear emergent world representation. Li's probe had been a small network of its own, because a straight line had failed. Nanda showed a straight line works once you ask the right question. The network stored each square as mine or theirs rather than black or white, and flipped its view every turn.
The same rule sits under both results, thirty-two years apart. Prediction rewards one thing only: making the next moment less surprising. Sometimes the cheapest way to be less surprised is to build the very thing that is making the data.
Predict the next what?
The loop says predict the next thing. It never says what counts as a thing, and that choice is the design decision. Let's take three choices in turn.
Predict the next pixel and you get a system that is very good at leaves. Leaves are hard to predict and they are also most of the pixels, so that is where the capacity goes (capacity is how much a model can hold. Spend it on one thing and it is not there for another). I call that the leaf problem, and chapter 4 is built around it.
Leaves are hard and they are most of the picture, so that is where the capacity goes. That is the leaf problem.
Target
pixel
What comes out
a renderer
Where the capacity went
the leaves
Predict the next word and you get grammar, plus a surprising amount of knowledge about the world. That is what the next word turns on. In 2019 Alec Radford and colleagues at OpenAI showed how far it goes at scale. Their paper was Language Models are Unsupervised Multitask Learners.
Their model was GPT-2. It was trained on nothing but next-word prediction over a large pile of web text. It came out able to answer questions, summarise, and translate a little, and nobody had trained it for any of those. OpenAI released it in stages, uneasy about what a fluent text machine might be used for.
The model learned those tasks because the next word sometimes depends on them. That is the clearest evidence we have that the target decides what gets built.
The capital of France is .
Fill the blank in your head, then press Reveal.
GPT-2 was trained only to predict the next word.
Sentence
1 of 5
What the next word depended on
try it yourself first
Predict the next state of some compact description and you get something a planner can use. Compact means a short list of numbers rather than a picture.
So the same loop gives three different machines, and all that changed is what was put on the right-hand side of the guess. Predict the next thing is a family of methods rather than one. Picking the target is the design decision that matters most, so an argument about prediction should say what is being predicted.
Being right is the same as being brief
There is a second way to look at this, four decades older than Elman's result, and it comes to the same thing. Shannon was a mathematician at Bell Labs, the research arm of the American telephone company, whose business was pushing more calls down the same wire. Suppose you have to send a message down a wire and you are charged by the symbol.
You and the receiver share a predictor. Before each symbol both of you run it and get a list of what is likely next, so you need only send enough to pick the right one. Confident and correct costs almost nothing, and confident and wrong costs a lot.
Roughly a coin toss. One or two bits. The price of a symbol is the probability you gave it.
Probability given
1 in 2
Cost to send
1.0 bits
Shannon made this exact in 1948. A symbol you gave a probability of one in two costs one bit (one bit is one yes or no answer, so eight bits can pick one option out of 256). One in four costs two, and one in a thousand costs about ten. That paper started information theory, and the word "bit" first appeared in print there.
Three years later he turned the argument on English itself. In Prediction and Entropy of Printed English he had people guess a passage one letter at a time. A language people can guess is cheap to send, and the experiment measured how cheap. That put a human being on the bottom row of Figure 3.10.
predict the next thing and check what actually arrived
Every symbol equally likely, so every symbol costs the same.
Bits per character
4.76
This message costs
257 bits
Shorter by
–
So the length of the shortest message you can write measures how well you predicted. That makes a better predictor a better compressor. Read backwards, whatever compresses your data the hardest has understood the most about it.
Marcus Hutter, a theorist of perfect learners in the abstract, took the reading literally. The Hutter Prize is a long-running cash prize for compressing a snapshot of Wikipedia. It runs on the argument that you cannot squeeze text much further without knowing what it means. Every winner to date has been, in effect, a language model.
Jürgen Schmidhuber went a step further in 2010. Schmidhuber ran a lab in Switzerland and had co-invented the LSTM, a memory network that was the workhorse of speech recognition for years. In 2018 he would co-write the paper this course is named after. His 2010 paper was Formal theory of creativity, fun, and intrinsic motivation.
He proposed rewarding a learner for compression progress. That is the discovery that it can now squeeze its data further than it could a moment ago. Getting better at compressing becomes a drive in its own right, a reason to go and look at new things.
Press Watch. The learner will look twelve times.
Step
0 of 12
Compressed size
100 of 100
Reward this step
0
The same reading says why the stowaways were never optional. Memorising is the costly way. A lookup table of everything that ever happened is huge, and useless on anything new.
The cheap option, the one that shortens the message, is to work out the rule that produced the data and keep that instead. The rule is the stowaway.
The table
- 1 0.00
- 2 0.31
- 3 0.59
- 4 0.81
- 5 0.96
- 6 1.02
- 7 0.99
- 8 0.87
- 9 0.67
- 10 0.41
- 11 0.12
- 12-0.18
- ... and 12 more
stores: 24 values
The rule
the rule: two numbers (illustrative)
next = a × last + b × the one before
stores: 2 numbers
The table is as big as the data. The rule is two numbers, whatever the data.
Values seen
24
The table stores
24 values
The rule stores
2 numbers
On new data
nothing new yet
Where it stops being enough
There are two limits.
Predicting well is not the same as being useful. A system can be excellent at continuing a recording and still hand you nothing you can act on. Saying what comes next is a different question from saying what would come next if you did something different. Only the second supports a decision. Li's network carries a board and was still only ever asked which moves were legal.
Answer
64
recorded next position
the recording is illustrative
The answer was already in the recording, one cell along. Nothing had to be worked out.
Question
what comes next
Where the answer comes from
the recording, one cell along
Moments to learn from
4
Being unsurprised is not the same as being right. A predictor that has only ever seen calm weather is not surprised by calm weather. That tells you nothing about what it would do in a storm.
Low error on what you happened to record is a weaker claim than it sounds, and it is the claim most headline numbers make. Let's see it with a ball instead of weather.
Inside the band it was shown. Anyone testing here would say it has the rule.
Where it really lands
69 m
Where the model says
71 m
Off by
4%
Neither limit undoes the argument. Prediction is enough to pull structure out of raw experience, and it is the cheapest signal anyone has found. The line from Shannon through Elman to Othello-GPT is people proving that with better instruments.
Everything here has assumed that what is being predicted is small: two numbers for a ball, and one word at a time for language. A camera hands you two million numbers, thirty times a second, and almost all of them are the leaves. That is the leaf problem arriving at full size, and it is where chapter 4 picks up.
Hopefully this chapter helped. There is a short quiz below if you want to check that the chain held together for you. If I have got something wrong, or you know a better source than the ones I used, the About page has a corrections section, and I would be glad to hear from you.
Try it
One photograph of a ball in flight. What can you not work out from it?
Pick one
1One photograph of a ball in flight. What can you not work out from it?
- a. Where the ball is
- b. How big the ball is
- c. Which way it is going and how fast
- d. What colour it is
Answer c. Which way it is going and how fast. Direction and speed are not in any single frame. They exist in the relationship between frames, which is the kind of thing a predictor has to build for itself.
2Why is next-thing prediction so much cheaper to train on than labelled data?
- a. The models are smaller
- b. The answer is already in the recording, so nobody has to write labels
- c. It needs less computing time
- d. It converges in fewer steps
Answer b. The answer is already in the recording, so nobody has to write labels. Every moment is the answer to the moment before. That turns any recording of anything into a training set, and removes the step where a person has to annotate a million examples.
3Jeffrey Elman's network was trained only to predict the next word, and its internal states sorted themselves into nouns, verbs, animate and inanimate. Why?
- a. Those categories were in the training labels
- b. The architecture had one unit per category
- c. A word's category is what its next word depends on, so the categories were the cheapest way to be less wrong
- d. It memorised the corpus
Answer c. A word's category is what its next word depends on, so the categories were the cheapest way to be less wrong. Nothing supplied the categories. Prediction rewards whatever makes the next thing less surprising, and for words that is grammatical category.
4A probe finds board state inside a network trained only on legal Othello moves. What turns that from a curiosity into a finding?
- a. The probe is very accurate
- b. The network was large
- c. Changing that internal board changes the moves the network then makes
- d. The board can be drawn as a picture
Answer c. Changing that internal board changes the moves the network then makes. Accuracy alone could be a coincidence in the numbers. Intervening on the representation and watching behaviour follow is what shows the network is using it.
5Under a good predictor, a symbol you were confident about and got right costs almost nothing to send. Why?
- a. It can be left out of the message
- b. The cost of a symbol is set by the probability you gave it
- c. Common symbols are stored in a table
- d. The receiver guesses it and does not need the message
Answer b. The cost of a symbol is set by the probability you gave it. Claude Shannon made the price exact. High probability means a short code, which is why a better predictor is a better compressor.
6What does the compression view say about memorising the training data?
- a. It is the best available strategy
- b. It is the expensive option: a lookup table of everything is enormous and useless on anything new
- c. It compresses better than any rule
- d. It is what all learning does
Answer b. It is the expensive option: a lookup table of everything is enormous and useless on anything new. The short message comes from finding the rule that generated the data and keeping that instead. Memorising is what compression penalises.
7Same loop, different target: predict the next pixel, or the next word, or the next compact state. What does that choice change?
- a. Nothing, the loop is what matters
- b. Only how long training takes
- c. What the system ends up good at, and what its capacity gets spent on
- d. Whether the method counts as self-supervised
Answer c. What the system ends up good at, and what its capacity gets spent on. Predict every pixel and most of the capacity goes to leaves, because that is where most of the pixels are. Picking the target is the design decision that matters most.
8A model has very low error predicting the next frame of a recording. What does that alone not tell you?
- a. That it saw the recording
- b. What it would predict if you acted differently, and how it behaves outside what it recorded
- c. That its error was measured correctly
- d. That the recording was long enough
Answer b. What it would predict if you acted differently, and how it behaves outside what it recorded. Being unsurprised by what you happened to record is a weaker claim than it sounds, and saying what comes next is not the same as saying what would come next under a different action.
Sources
- A Mathematical Theory of CommunicationShannon, 1948Where the price of a symbol becomes the probability you gave it. Everything about prediction and compression being one job starts here.
- Prediction and Entropy of Printed EnglishShannon, 1951Claude Shannon sat people down and had them guess English one letter at a time. The bottom row of Figure 3.10 is roughly what he measured.
- Finding Structure in TimeElman, 1990Train a small network to predict the next word, then look inside: nouns, verbs, animate and inanimate, none of it asked for.
- Emergent World RepresentationsLi et al., 2022A network given nothing but legal Othello moves, with the board found inside it and causally manipulated.
- Othello-GPT has a linear emergent world representationNanda et al., 2023The follow-up that sharpened what the probe was actually reading.
- Language Models are Unsupervised Multitask LearnersRadford et al., 2019Next-word prediction, scaled, turning into capabilities nobody trained for. The strongest evidence that the target choice is the design decision.
- The Hutter PrizeHutter, ongoingA cash prize for compressing a snapshot of Wikipedia, run on the argument that you cannot squeeze text further without modelling what it means.
- Representation Learning: A Review and New PerspectivesBengio, Courville & Vincent, 2013The survey that laid out what a good learned representation is for, written before prediction had finished winning the argument.
- Formal theory of creativity, fun, and intrinsic motivationSchmidhuber, 2010Compression progress as a drive in its own right: not just a way to measure a model, but a reason to go and look at something.
This chapter is free and open source. A star helps the next person find it.
Star on GitHub21