cosmic college · ai 101
ai 101 · thinking in math
Artificial intelligence, taught the way a math class would teach it. It starts with the story of how decades of patient ideas became tools we use every day. Then bits, functions, vectors, and learning by small corrections, with simple numbers you can follow by hand.
math that talks back
You type a question into a chat window, and a few seconds later a paragraph appears that reads as if someone thought about it. It feels like magic. Underneath, it is arithmetic: multiplication and addition, done an enormous number of times, very quickly, by machines that only understand on and off.
This lesson treats AI the way a math class would. No code and no walls of jargon. We start with bits, build up to functions, vectors, and learning by small corrections, then see how stacking simple pieces gives you something that can write a poem or translate a menu. First comes a short history, because every idea here took decades of patient work to arrive.
If you can multiply two decimals and follow a recipe, you can follow every step below.
the story · from turing to chat
In 1950 the British mathematician Alan Turing published a paper that opens with a simple line: “I propose to consider the question, ‘Can machines think?’” Rather than argue over definitions, he proposed a game. If a person chatting by text could not reliably tell a machine from a human, perhaps the question would answer itself. He called it the imitation game.
Six years later, in the summer of 1956, a small group gathered at Dartmouth College for a workshop proposed by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. Their proposal gave the field its name, artificial intelligence, and suggested that a carefully chosen group working together for one summer could make real progress. The full answer has taken a good deal longer.
In 1958 Frank Rosenblatt, a psychologist at the Cornell Aeronautical Laboratory, unveiled the perceptron, a program that learned to tell cards marked on the left from cards marked on the right by adjusting its weights after each mistake. A hardware version followed, with 400 light sensors for eyes. Newspapers imagined thinking machines within a few years. In 1969 Minsky and Seymour Papert published a careful book showing what a single layer of weights cannot learn, and as promises ran ahead of results, funding cooled through the 1970s. A second chill followed the expert-system boom of the 1980s. People came to call these stretches the AI winters.
A few researchers kept working through the cold. In 1986 David Rumelhart, Geoffrey Hinton, and Ronald Williams showed how a method called backpropagation could train networks with hidden layers, and Hinton, Yann LeCun, and Yoshua Bengio later shared the 2018 Turing Award for their work on deep neural networks. Meanwhile, on May 11, 1997, IBM’s Deep Blue won the deciding game of a six-game rematch against world chess champion Garry Kasparov. Deep Blue did not learn the way modern models do. It searched about 200 million chess positions a second, guided by rules its builders tuned by hand, and it still felt like a line had been crossed.
Learning returned in 2012. Alex Krizhevsky, Ilya Sutskever, and Hinton at the University of Toronto trained a deep network, later called AlexNet, on two gaming graphics cards and entered the ImageNet challenge. Its top-5 error was 15.3 percent; the next best entry scored 26.2. In 2017 eight researchers at Google published “Attention Is All You Need,” introducing the transformer. On November 30, 2022, OpenAI released ChatGPT as a free research preview, and five days later it had a million users. Seventy-two years after Turing asked his question, millions of people were asking a machine questions of their own.
timeline · from turing to language models
Grey bands mark the two AI winters, when funding and interest dropped. Dates are the commonly cited years for each milestone.
The big ideas are old. The recent change is scale. Patient research plus abundant data and compute turned math from the 1950s and 1980s into everyday tools.
what computing is
A computer stores everything as bits. A bit is a switch with two states, written 0 and 1. One bit tells you very little. Line bits up and the options multiply: every extra bit doubles the number of patterns you can make.
worked example · counting patterns
patterns = 2number of bits
1 bit gives 2 patterns. 3 bits give 2 × 2 × 2 = 8. 8 bits, one byte, give 256.
In the long-standing ASCII standard, the capital letter A is the number 65, stored as the byte 01000001. Text, photos, and sound are all agreed-upon patterns like this one.
Bits become useful when simple rules combine them. Three rules do most of the work. AND outputs 1 only if both inputs are 1. OR outputs 1 if either input is 1. NOT flips a bit. Wire a few of these together and you can add:
| a | b | sum bit (one or the other) | carry bit (a AND b) |
|---|---|---|---|
| 0 | 0 | 0 | 0 |
| 0 | 1 | 1 | 0 |
| 1 | 0 | 1 | 0 |
| 1 | 1 | 0 | 1 |
Read the last row as 1 + 1 = 10 in binary, which is two. Chain these little adders and you get a circuit that adds any pair of numbers. Chain enough circuits and you get a processor.
That is the central idea of computing: complicated behavior built from very simple rules, repeated and connected. Every program is a function. It takes an input, follows fixed steps, and returns an output.
A computer does a small number of very simple things, extremely fast. The power comes from how many times it does them and how they are arranged.
a model is a function
In math class, a function is a rule that turns an input into an output. The function f(x) = 2x + 1 takes 3 and returns 7. Give it the same input and you always get the same output.
An AI model is a function too. The input might be a photo, a sentence, or a column of sales figures. The output might be a label, the next word, or a forecast. The difference is how the rule gets written. Nobody types it in by hand. The rule contains adjustable numbers, called weights or parameters, and those numbers are set by learning from examples.
same shape · input, rule, output
Both boxes are functions. In the top row a person wrote the rule. In the bottom row the rule is a huge arrangement of multiplications and additions whose numbers were learned from examples.
Take a tiny model that guesses how long a bike ride will take: predicted minutes = w × kilometers. The single weight w is the model's entire knowledge. If w = 4, a 2 km ride is predicted at 8 minutes. Change w and the model changes its mind. Large models work the same way, with billions of weights in place of one.
Model = function + adjustable numbers. Training = choosing the numbers. Using the model = plugging in new inputs.
numbers as vectors
Models only handle numbers, so everything has to become numbers first. The trick is to describe each thing with a list of numbers, called a vector. Each position in the list measures one feature.
Here is a toy version with two made-up features, how alive something is and how much it has an engine, each scored from 0 to 1:
- cat = (0.9, 0.1)
- dog = (0.8, 0.2)
- car = (0.1, 0.9)
Draw each list as an arrow from the origin and similar things point in similar directions. The number that measures this is the dot product: multiply matching positions, then add.
worked example · the dot product
a · b = (a1 × b1) + (a2 × b2)
cat · dog = (0.9 × 0.8) + (0.1 × 0.2) = 0.72 + 0.02 = 0.74
cat · car = (0.9 × 0.1) + (0.1 × 0.9) = 0.09 + 0.09 = 0.18
The bigger score says cat and dog agree more than cat and car do. When arrows have different lengths, dividing the dot product by both lengths gives cosine similarity, a pure measure of direction.
vectors · direction means similarity
Cat and dog sit almost on top of each other, so their dot product is high. Car points toward the engine axis, so its score with either animal is low. Real models use hundreds or thousands of axes, and the geometry works the same way.
Real models learn their own features instead of using ones we name, and each vector has hundreds or thousands of positions instead of two. The geometry still holds: related ideas land near each other. A well-known 2013 result showed word vectors where king − man + woman lands close to queen, a sign that directions in the space can carry meaning.
Keep the dot product in mind. It is the most repeated operation in modern AI, and it comes back twice more in this lesson.
learning by small steps
Back to the bike model, minutes = w × kilometers. Suppose we timed one real ride: 2 km took 6 minutes. We start with a poor guess, w = 1, and let the model learn.
First we need a score for how wrong the model is. A common choice is squared error: (prediction − actual)². Squaring keeps the score positive and punishes big misses more than small ones. That score is called the loss.
Then we ask a simple question: if we nudge w up a little, does the loss go up or down, and how fast? That rate is the slope, also called the gradient. We take a step in the downhill direction, sized by a small number called the learning rate. Then we repeat. That loop is gradient descent.
worked example · gradient descent by hand
loss = (2w − 6)²
slope = 2 × (2w − 6) × 2
new w = w − 0.1 × slope
- start: w = 1. Prediction 2 minutes, error −4, loss 16. Slope = 2 × (−4) × 2 = −16.
- step 1: w = 1 − 0.1 × (−16) = 2.6. Prediction 5.2, error −0.8, loss 0.64.
- step 2: slope = 2 × (−0.8) × 2 = −3.2, so w = 2.6 + 0.32 = 2.92. Prediction 5.84, loss 0.0256.
- step 3: slope = −0.64, so w = 2.984. Loss is about 0.001.
Three steps took the loss from 16 to about 0.001, and w settled toward 3 minutes per kilometer, exactly what the ride showed. Nobody told the model the answer. It followed the slope.
You can check the slope without calculus. At w = 1 the loss is 16. At w = 1.01 the prediction is 2.02, the error is −3.98, and the loss is 15.84. A nudge of 0.01 lowered the loss by about 0.16, which is 16 times the nudge. That is what a slope of −16 means.
gradient descent · rolling downhill
The loss makes a bowl. The dashed line is the slope at the start: steep and pointing downhill to the right. Each step moves w toward the bottom, and the steps shrink as the bowl flattens.
Real training does exactly this with billions of weights at once. The method that computes all those slopes efficiently, layer by layer, is called backpropagation. Every weight gets its own tiny nudge on every step, and the loss drifts down over millions of steps across huge piles of examples.
Learning = measure the error, find the downhill direction, take a small step, repeat. Step size matters: in our example, a learning rate of 0.3 would jump past 3 to w = 5.8, with a loss of 31.36, and bounce farther away each time.
stacked simple functions
One weight can learn one straight-line relationship. The world is rarely a straight line. Neural networks get their flexibility by stacking many small functions, each one easy to understand on its own.
The basic unit is a neuron, a loose borrowing from biology. It does three things: multiply each input by a weight, add the results along with one extra number called a bias, then pass the total through a simple bend. A common bend is ReLU: keep positive numbers as they are and turn negative numbers into 0.
worked example · one neuron
output = max(0, w1x1 + w2x2 + b)
Weights (0.5, −1), bias 4. Inputs (2, 3): 0.5 × 2 + (−1) × 3 + 4 = 1 − 3 + 4 = 2. Positive, so the output is 2.
Inputs (2, 6): 1 − 6 + 4 = −1. Negative, so the bend makes the output 0.
The weighted sum is a dot product, the same one from the cat and dog example, plus the bias.
Why the bend? Without it, stacking adds nothing new. If one layer doubles a number and the next triples it, the pair simply multiplies by 6, one straight line again. The bend lets each layer fold the problem into pieces, and enough pieces together can trace almost any shape. The early single-layer perceptron hit exactly this wall: it could not learn XOR, the rule "one or the other, not both".
a small network · 3 → 4 → 4 → 2
Information flows left to right. The highlighted path is one of many routes a signal can take. Training adjusts all 46 numbers together, using the same downhill steps as the bike model.
That small drawing already has 46 adjustable numbers. Large language models follow the same pattern with many more layers and billions of parameters. Training sets every one of them with the downhill-step method from the last section.
A neural network is simple functions, stacked. Each piece is a dot product and a bend. Depth and width give it flexibility.
why scale mattered
Most of the ideas above existed by the late 1980s. What changed afterward was scale, along three dials that grew together:
- data: the web put enormous amounts of text and images within reach, and projects such as ImageNet (2009) labeled millions of photos for training and testing.
- parameters: bigger networks can hold more patterns and subtler ones.
- compute: graphics chips (GPUs), built to shade millions of pixels in parallel, turned out to be excellent at the same multiply-and-add work that neural networks need.
In 2020 researchers measured something striking. As model size, data, and compute grew together, the loss fell smoothly and predictably, following a simple curve called a power law. That made progress something you could plan. In 2022 a follow-up study found that many large models had been trained on too little data for their size, and that data and parameters should grow roughly in step, at around 20 training tokens per parameter.
worked example · sizing a training run
training compute ≈ 6 × parameters × training tokens
This is a widely used approximation. Assumption: a 1 billion parameter model trained on 20 billion tokens, following the 20-per-parameter guideline.
Compute ≈ 6 × 109 × 2 × 1010 = 1.2 × 1020 arithmetic operations. Assumption: a chip sustaining 1014 useful operations per second. Time = 1.2 × 1020 ÷ 1014 = 1.2 × 106 seconds, about 14 days on one chip.
Make the model 10 times bigger and feed it 10 times more tokens, and the compute grows 100 times.
This is where the track bends toward hardware. Once progress follows compute, chips, power, cooling, and datacenters become part of the story. That is the ground 201 covers.
Three dials: data, parameters, compute. Turn them together and results improve predictably. That predictability is why compute became a planning question.
tokens and next words
Language models read text as tokens: common words, pieces of longer words, punctuation, and spaces. A word like "unbelievable" might become three tokens, such as un · believ · able. Each token has an ID number from a fixed vocabulary, typically tens of thousands to a couple hundred thousand entries, and each ID maps to a learned vector, just like cat and dog earlier.
The training task is almost disarmingly simple: given the tokens so far, predict the next one. The model produces a score for every token in its vocabulary, and those scores are turned into probabilities that add up to 1.
next-token prediction · one step
The model does not store one answer. It assigns a probability to every token, picks one (often a likely one, with a little randomness), appends it, and runs again.
The loss is how surprised the model was by the real next token. If the text actually said "mat" and the model gave mat 0.41, that is a little surprise. If it gave mat 0.01, that is a lot. Gradient descent nudges billions of weights to be a bit less surprised, across trillions of tokens of text.
To write, the model repeats one move: predict a distribution, pick a token, add it to the text, predict again. A paragraph is a few hundred of those loops. To get good at guessing the next word across all kinds of writing, a model has to pick up grammar, facts, and patterns of reasoning along the way. That is why such a plain objective turned out to be so capable.
Transformers, introduced in 2017, add one more idea to the stacked layers: attention. At each layer, every token scores the tokens before it with a dot product and blends in information from the ones that score highest. That is how a word like "it" can find the noun it refers to, many words back.
A language model is a next-token function. It has no lookup table of answers. It generates from patterns stored in its weights, which is why it can be fluent and still wrong. Check the facts that matter.
closing rules of thumb
- Computing is simple rules on bits, repeated and connected.
- A model is a function with adjustable numbers inside. Training picks the numbers.
- Vectors turn things into arrows; dot products measure how much two arrows agree.
- Learning is gradient descent: measure the error, step downhill, repeat.
- Neural networks stack simple pieces, each a dot product and a bend.
- Data, parameters, and compute grew together, and results improved predictably.
- A language model predicts the next token, one loop at a time. It is fluent by design, so check the facts that matter.
- The machinery is math, all the way down, and that is the wonder of it.
Next in this track, 201 · under the hood follows the math onto real hardware and software: chips, memory, power, and what it costs to train and run a model. 301 · datacenters in space asks what changes when that hardware leaves the planet. Both will appear in cosmic college when the drafts are ready.