Generative AI Engineer Series · Article 01 · Phase 1: Foundations
What Is Generative AI?
From machine learning to large language models. What it means for a machine to learn, how a next-word predictor turns into an assistant that writes essays and code, where it goes wrong, and where this series will take you.
- One small photo
- 196,608numbers a computer sees in a 256 × 256 leaf image
- GPT-3 (2020)
- 175Bparameters, each one an adjustable number
- Pre-training data
- Trillionsof tokens, labelled for free by the next word
- Training stages
- 3from raw text predictor to helpful assistant
§ 01 · HookIt Isn't Magic
Open ChatGPT, Claude, or Gemini and type: "Explain compound interest to a twelve-year-old using pocket money in rupees." A few seconds later you have a friendly, correct, well-organised explanation, complete with a worked example about saving ₹100 a month. Ask it to turn that explanation into a short poem, and it does. Ask for Python code that calculates the final amount, and you get that too.
You've probably met the rest of the family as well. Midjourney turns a single sentence into a photorealistic image, GitHub Copilot finishes the function you were halfway through writing, and some voice assistants now sound uncannily human.
It's tempting to file all this under "magic" and move on. It isn't magic, though. Every one of those answers comes out of a chain of ideas that has been building for decades, and a motivated beginner can understand every link in that chain without advanced maths or a computer science degree.
That's what this article is for.
This is Article #01 of the Generative AI Engineer Series, a structured path that starts with the fundamentals and ends with you building and deploying a full GenAI application from scratch. Every later article assumes you roughly know what machine learning is, what a neural network does, what it means for a model to be "trained", and what makes a model generative. This article builds all of that from zero, using arithmetic, everyday analogies, and two short optional code snippets for readers who like to see things run.
By the end, you'll be able to answer these questions with confidence:
- What is the actual difference between AI, machine learning, deep learning, and a large language model?
- What does it mean for a machine to "learn" anything at all?
- Why did we need neural networks, and what makes them "deep"?
- How can something as simple as predicting the next word produce a system that writes essays and code?
- What makes a model generative, and what kinds of generative AI exist beyond chatbots?
- How does a raw text predictor become a polite, helpful assistant?
- How do foundation models change the way AI products get built, and when is traditional ML still the better tool?
- Why do these models sometimes make things up with complete confidence, and what do engineers do about it?
- Why do tricks like "give it a few examples" or "ask it to think step by step" actually work?
- Where is all this used in practice, and what will you be able to build by the end of the series?
§ 02 · The Big PictureFour Nested Ideas
People use AI, machine learning, deep learning, and LLM almost interchangeably, but they mean different things. The easiest way to see how they relate is to picture a set of Russian nesting dolls, each one sitting inside the one before it.
← swipe to see the whole figure →
Artificial Intelligence (AI) is the outermost doll and the broadest idea. It covers any technique that makes a machine do something we'd call intelligent if a person did it: old-fashioned chess programs that follow hand-written rules, the route planner in your maps app, modern chatbots. AI describes a goal rather than a particular method.
Machine Learning (ML) sits inside AI. It's one particular approach to building intelligent behaviour, where the system learns the rules from examples instead of a programmer writing them by hand.
Deep Learning (DL) sits inside machine learning. It's a family of ML methods built on multi-layer neural networks, and it turned out to be spectacularly good at messy, unstructured data like photos, audio, and text. (Structured data fits neatly into a spreadsheet: rows of customers with columns for age, city, and income. Unstructured data is everything that doesn't, such as a photograph, a voice note, or a paragraph of writing.)
Large Language Models (LLMs) sit inside deep learning. They're deep learning models trained on huge amounts of text, and they power ChatGPT, Claude, Gemini, and Llama. They're also the best-known members of a broader category called Generative AI: models that create new content instead of only labelling content that already exists. §09 pins down what "generative" means, and §11 covers the other kinds of generative AI, for images, audio, code, and video.
We'll open the dolls from the outside in, picking up the ideas we need at each layer.
The shift that defines machine learning: from writing rules to learning them
Before machine learning, getting a computer to do something meant writing explicit instructions. Say you want to build an email spam filter the traditional way. You might write:
IF the email contains "lottery" AND "claim your prize" → mark as spam IF the sender is not in the contact list AND the email contains a link → mark as spam
This works for about a week. Then spammers start writing "l0ttery" with a zero, and a genuine email from your bank that mentions "claim" gets blocked. You add more rules, then exceptions to the rules, then exceptions to the exceptions. The rulebook turns into an unmaintainable mess, and it's always one step behind the spammers.
Machine learning flips the process around:
- Traditional programming: you provide rules + data, and the computer produces answers.
- Machine learning: you provide data + answers, and the computer produces the rules.
For the spam filter, you'd collect 100,000 emails that people have already marked as "spam" or "not spam", hand them to a learning algorithm, and let it work out which patterns matter and by how much. When spammers change tactics, you don't rewrite any rules. You collect fresh examples and retrain.
← swipe to see the whole figure →
Think about how a child learns what a dog is. Nobody hands them a definition, and "four legs, fur, a tail" would describe a cat just as well anyway. They see lots of dogs and lots of things that aren't dogs, get gently corrected when they point at a goat and say "doggy", and gradually work out the pattern. Machine learning is a mathematical version of the same process: it learns from examples and corrections rather than from definitions.
§ 03 · Machine LearningLearning Patterns From Examples
A worked example: sorting mangoes
Ramesh runs a mango packing unit near Lucknow. Thousands of mangoes roll down his conveyor belt every day, and each one has to be sorted as ripe (ship to local markets now) or unripe (store for a few days, or send to distant cities). At the moment two workers do this by hand, all day long. Ramesh wants to automate it.
He fits two cheap sensors over the belt:
- A firmness sensor that gently presses each mango and gives a score from 0 (very soft) to 10 (rock hard).
- A colour sensor that gives a score from 0 (completely green) to 10 (fully yellow-orange).
For one week his workers keep sorting by hand as usual, but now each mango's two sensor readings get recorded next to the worker's decision. By the end of the week Ramesh has a table of 2,000 mangoes, each with a firmness score, a colour score, and a label: ripe or unripe.
That small example already contains the core vocabulary of machine learning:
- Features are the measurable inputs. Here, firmness and colour.
- The label is the correct answer we want to predict. Here, ripe or unripe.
- Training data (or the dataset) is the collection of examples where we know both the features and the label.
- The model is the thing that learns the relationship between features and label, so it can predict the label for new examples where we only have the features.
Plot every mango on a graph, with firmness along one axis and colour along the other, and a clear picture shows up. Ripe mangoes cluster in the soft-and-yellow corner and unripe ones in the hard-and-green corner. There's some overlap in the middle, but the pattern is obvious.
← swipe to see the whole figure →
Training the model means finding the line that best separates the two groups. That line is called the decision boundary. Once the model has it, sorting a new mango is easy: measure it, place it on the graph, and check which side of the line it lands on.
There's a bonus, too. A mango far from the line, deep in the ripe cluster, is almost certainly ripe, while one sitting right on the line could honestly go either way. So most models output a probability rather than a bare label, something like "93% ripe" or "52% ripe". That's useful in practice: Ramesh can let the machine handle the confident cases and send only the uncertain ones to a worker for a second look.
Classification vs regression
Ramesh has a classification problem, because the answer is one of a fixed set of categories (ripe or unripe). Spam vs not-spam is classification, and so is working out which of ten digits a handwritten number is, or whether a transaction is fraudulent.
The other main kind of problem is regression, where the answer is a number on a continuous scale. Predicting the price of a 2BHK flat in Pune from its area, floor, and locality is regression. So is forecasting tomorrow's temperature or next month's sales.
Both are supervised learning, meaning learning from examples that come with the correct answers attached, like a student working through a textbook with an answer key. (There's also unsupervised learning, where the data has no labels and the goal is to find structure on its own, for example grouping customers into segments nobody defined in advance. We won't need it in this article.)
What "training" actually looks like
How does the model find the best line? It doesn't know it at the start. It begins with a random line that sorts the mangoes terribly, then repeats a simple loop:
- Check how wrong the current line is: how many mangoes land on the wrong side, and how badly.
- Nudge the line slightly in whichever direction reduces that wrongness.
- Repeat, hundreds or thousands of times.
Each nudge is small, but over many rounds the line settles into the position that makes the fewest mistakes. That measure of "how wrong" is called the loss, and minimising it is the heart of all machine learning. We'll look inside this loop properly in §05, because the loop that trains ChatGPT is the same one, run at an unimaginably larger scale.
Learning vs memorising: generalisation and overfitting
This is the most important idea in machine learning, and the one beginners most often miss: the goal is to do well on data the model has never seen. Doing well on the training data is only a means to that end. The ability to handle unseen data is called generalisation.
To measure it, we always hold back part of the labelled data (say 20% of Ramesh's mangoes) and never show it to the model during training. This held-back part is the test set. After training, we check how well the model sorts these unseen mangoes, and that's the score that actually counts.
Why does this matter so much? Picture two students preparing for an exam. One memorises the last five years of question papers word for word, answers and all. The other works at understanding the underlying concepts. On a practice test built from those old papers, the first student scores 100%. In the real exam, where the questions are reworded, the first student falls apart and the second does fine.
A model can fall into the same trap. If it's flexible enough, it can draw a wildly wiggly boundary that twists around every training mango, including the handful a tired worker mislabelled. That model scores perfectly on the training data and badly on new mangoes. This is overfitting. The opposite failure, a model too simple to capture the real pattern, is underfitting.
← swipe to see the whole figure →
When the world gets messier
Ramesh's problem is a friendly one. Real problems get harder in three ways at once:
- The boundary isn't straight. Very soft, very yellow mangoes might be overripe, so ripeness follows a curve rather than a line.
- There are more features. Weight, smell, variety (an Alphonso ripens differently from a Dasheri), days since harvest. You might have ten, fifty, or hundreds of inputs instead of two.
- There are more classes. Unripe, ripe, overripe, damaged, wrong variety.
Which leads to the idea that runs through the rest of this article:
The more complex the relationship between inputs and outputs, the more flexible the model has to be to learn it, and the more data you need to learn it reliably without overfitting.
Straight lines are fine for mangoes. They're hopeless for what comes next.
§ 04 · The WallWhere Classic Machine Learning Hits a Wall
The mango problem has two very convenient properties. There are only a few inputs, and each input obviously means something: firmness matters, colour matters. Two kinds of real-world problems break both properties at once.
Problem 1: images
A farmer in Maharashtra photographs a tomato leaf with her phone. We want a model that says whether the leaf is healthy, has early blight, or has leaf curl virus, so she can act before the disease spreads across the field.
To a computer, an image is a grid of tiny coloured squares called pixels. Each pixel is stored as three numbers for how much red, green, and blue it contains, each from 0 to 255. So a modest 256 × 256 photo contains 256 × 256 × 3 = 196,608 numbers.
(firmness, colour)
256 × 256 leaf photo
review (768 per word)
The sheer count isn't even the real problem. The real problem is that no individual number means anything on its own. Knowing that the pixel at row 120, column 87 is slightly brown tells you nothing. The disease lives in the pattern (ring-shaped brown spots with yellow halos), and that pattern can show up anywhere on the leaf, at any size, in bright sun or cloudy shade, with the leaf tilted at any angle. The relationship between those 196,608 raw numbers and the correct label is hugely complicated, and a straight line, or even a gentle curve, has no chance of capturing it.
← swipe to see the whole figure →
Problem 2: text
Now take a food-delivery review:
"Absolutely loved waiting ninety minutes for cold biryani. Ten stars, would starve again."
Word by word, it screams positive: "loved", "ten stars", "would … again". The actual meaning is furious. People spot the sarcasm instantly. A model that looks at one word at a time doesn't.
Text has an extra complication that images don't: it isn't numbers to begin with. Before a model can process a word, we need some way to turn it into numbers. The naive approach is to give each word an ID, so "good" is 4,031 and "excellent" is 9. That's useless, because the IDs say nothing about meaning. "Excellent" isn't a tiny bit of "good" just because 9 is smaller than 4,031.
The modern solution is embeddings. Each word (or piece of a word) is represented as a list of hundreds of numbers, learned so that words with similar meanings end up with similar lists. Embeddings get a full article of their own (Article #04). For now, just notice the scale. A 300-word review, with each word represented by 768 numbers, becomes 230,400 numbers. On top of that, order matters ("dog bites man" vs "man bites dog"), context changes meaning ("not bad at all" is praise), and tone can flip everything.
The real bottleneck: feature engineering
For decades, researchers handled images and text with a workaround called feature engineering. A human expert hand-designed a small number of meaningful measurements, and a classic ML model learned from those instead of from the raw data.
For the tomato leaves, an expert might write code to compute "percentage of brown pixels", "number of roughly circular spots", and "average spot size". For reviews, you might count how many words appear on a list of positive words versus a list of negative ones.
This sort of worked, but it had serious weaknesses:
- It was slow and expensive. Every new problem needed months of expert effort to design good features.
- It was brittle. Features that worked on photos taken in sunlight failed on photos taken indoors.
- It was capped by human imagination. The model could only use patterns someone thought to measure. Nobody has written a hand-made "sarcasm detector" feature that really works, because sarcasm is exactly the sort of pattern that resists simple rules.
What we actually wanted was a model flexible enough to learn arbitrarily complicated relationships directly from raw pixels and raw words, one that would discover useful features by itself. That's what deep learning delivers.
§ 05 · Deep LearningNeural Networks
The artificial neuron
The building block of deep learning is the artificial neuron, which is far simpler than the name suggests. A neuron does three things:
- It takes in several numbers.
- It multiplies each one by a weight (a number for how much that input matters), adds them all up, and adds one more adjustable number called a bias.
- It passes the total through an activation function, which decides how strongly the neuron "fires" in response.
Imagine a bank loan officer deciding whether to approve an application. Monthly income matters a lot, so it gets a large positive weight. The number of existing loans counts against approval, so it gets a negative weight. The applicant's favourite colour is irrelevant, so its weight is roughly zero. The officer mentally adds up the weighted evidence and makes a call. A neuron does the same thing with numbers.
In fact, a single neuron is basically our mango model: weigh the inputs, add them up, draw a line. The name was loosely inspired by biological brain cells, but don't push the comparison too far. An artificial neuron is a small piece of arithmetic, and it doesn't model how a real brain cell works.
Layers
A neural network arranges lots of neurons into layers:
- The input layer receives the raw numbers, such as pixel values or word embeddings.
- One or more hidden layers sit in the middle. Each neuron in a hidden layer takes the outputs of all the neurons in the layer before it.
- The output layer produces the final answer, usually one number per possible class, converted into probabilities ("78% early blight, 15% healthy, 7% leaf curl").
A network with many hidden layers is called deep, and that's all "deep learning" means: learning with neural networks that have many layers.
← swipe to see the whole figure →
Why the "bend" matters: non-linearity
This next point is subtle, and a lot rests on it. If every neuron only computed a weighted sum, stacking layers would be pointless. A weighted sum of weighted sums is still just a weighted sum; mathematically, any stack of purely linear steps collapses into a single linear step. You could stack a thousand layers and still only draw straight lines.
The activation function is what breaks this, by introducing a bend. The classic choice, ReLU (Rectified Linear Unit), is almost comically simple: if the number is negative, output zero, otherwise pass it through unchanged. That tiny kink is enough. (Modern LLMs mostly use smoother cousins of ReLU, but the job they do is the same.)
Think about folding a sheet of paper. One straight fold gives you one crease, but many folds at different angles and positions can produce astonishingly intricate shapes. That's origami. Each neuron's activation is a small fold, and a network with millions of them can bend its decision boundary into almost any shape the data calls for. Mathematicians have proved that a large enough network can approximate essentially any smooth input-output relationship, which is the flexibility the leaf and review problems needed.
Depth means learning features automatically
Why use many layers instead of one enormous layer? Because each layer builds on the one before it, which gives you a natural hierarchy of features.
When researchers look inside networks trained on images like our tomato leaves, they typically find something like this:
- Early layers respond to simple things: edges, colour contrasts, small textures.
- Middle layers combine those into shapes: spots, rings, the branching pattern of veins.
- Later layers combine shapes into high-level concepts such as "blight-like lesion" or "curled leaf margin".
Nobody programmed any of these detectors. They emerge during training because they turn out to help reduce errors. That's the answer to the feature-engineering bottleneck from §04: the network engineers its own features, layer by layer.
Text works the same way. Lower layers pick up word pieces and local phrases, and higher layers capture grammar, meaning, tone, and yes, sarcasm.
← swipe to see the whole figure →
How a network learns: loss, gradient descent, and backpropagation
Every weight and bias in the network is a parameter, an adjustable number. A freshly created network has random parameters, so its predictions are garbage. Training means tuning millions (or billions) of these numbers until the predictions get good. It uses the same loop we saw with the mangoes, now with proper names:
1. Forward pass. Show the network one example, say a leaf photo we know has early blight. The numbers flow through the layers and a prediction comes out, perhaps "40% early blight".
2. Loss. Compute a single number that measures how wrong that was. The true answer was early blight and the network only gave it 40%, so the loss is fairly high. If it had said 99%, the loss would be close to zero.
3. Gradient descent. Now adjust the parameters to bring the loss down. Imagine you're a hiker on a hillside in thick fog, trying to reach the lowest point of the valley. You can't see anything, but you can feel which way the ground slopes under your feet. So you take a small step downhill, feel the slope again, take another step, and keep going until you reach the bottom.
← swipe to see the whole figure →
Training does the same thing. The hillside is the loss, and your position on it is set by the values of all the network's parameters. For each parameter we calculate the slope, called the gradient, which tells us which direction to nudge that parameter to reduce the loss. Then we nudge all of them a tiny bit downhill. The size of each step is the learning rate. Too large and you overshoot the valley; too small and training takes forever.
4. Backpropagation. With billions of parameters, how do we calculate every gradient efficiently? With backpropagation. Starting from the error at the output, the calculation works backwards through the network, layer by layer, figuring out how much each parameter contributed to the mistake.
Think of a restaurant tracing a customer's complaint about a bad dish. The manager works backwards through plating, cooking, prep, and the ingredient supplier, and adjusts each station in proportion to how much it contributed to the problem. Backpropagation is that blame-assignment process, done with calculus, for every parameter at once.
Repeat the loop over millions of examples, usually passing through the whole dataset several times (each full pass is an epoch), and the random network gradually becomes an accurate one.
Parameters vs neurons: a common confusion
Model sizes are quoted in parameters, not neurons, and the two numbers are very different. A layer of 1,000 neurons, each connected to all 1,000 neurons in the previous layer, has 1,000 × 1,000 = one million weights but only 1,000 neurons. GPT-3, released in 2020, had 175 billion parameters.
You'll sometimes see these figures compared with the roughly 86 billion neurons in a human brain. That comparison mixes up units. A model's parameter is closer to a connection between brain cells (a synapse) than to a brain cell, and the brain has something like a hundred trillion synapses, possibly more. The comparison is fun, but it doesn't hold up once you look at it closely.
Why deep learning took off when it did
Neural networks aren't new. The core ideas are decades old, and backpropagation was popularised in the 1980s. So why did deep learning only take off in the 2010s? Because three ingredients finally showed up at the same time:
- Data. Large labelled datasets appeared, most famously ImageNet, with millions of hand-labelled images.
- Compute. GPUs (graphics processing units), originally built to render video games, turned out to be perfect for the huge volume of simple multiply-and-add arithmetic neural networks need, because they run thousands of calculations in parallel.
- Techniques. Better activation functions, better ways to initialise weights, and better tricks for preventing overfitting.
In 2012 a deep network called AlexNet won the ImageNet image-recognition competition by a wide margin over traditional methods. That result is usually treated as the start of the modern deep learning era.
§ 06 · Large Language ModelsWhat the Name Actually Means
We now have everything we need to take the name apart.
"Large" refers to the number of parameters. Today's leading models have billions to hundreds of billions of them, a few have passed a trillion, and for many frontier models the exact count isn't public. There's no official cut-off for "large". Models with a few billion parameters or fewer are increasingly called small language models, and they can run on a laptop or even a phone.
"Language model" is the more interesting part. A language model has learned to answer one question extremely well:
Given some text, what comes next?
Next-word prediction is classification (a very big one)
Remember the promise from §03? This is where it pays off. Look at next-word prediction through the lens of the mango problem:
| Mango sorting | Language modelling | |
|---|---|---|
| Input | Firmness and colour | A sequence of text |
| Output | One class from a fixed list | One token from a fixed list |
| Number of classes | 2 (ripe, unripe) | ~50,000–200,000+ |
| Type of problem | Classification | Classification |
It's the same kind of problem. The input is much richer and the list of possible answers is enormous, but that's the only difference.
One small detail: language models don't actually predict whole words. They predict tokens, which are frequently occurring chunks of text. A common word like "the" is a single token, while a rarer or longer word may be split into two or three pieces. The fixed list of all the tokens a model knows is its vocabulary. Article #03 covers tokens in depth. For now you can think "token ≈ word" without losing anything important.
So given the text "The monsoon arrived early, so the farmers started", a language model outputs a probability for every single token in its vocabulary:
← swipe to see the whole figure →
A tiny language model you can run (optional)
To make this concrete, here's the simplest possible language model. It reads a small piece of text and counts which word follows which. Then, for any word, it turns those counts into probabilities.
from collections import Counter, defaultdict import random corpus = """the chai is hot . the chai is sweet . the coffee is hot . the rain is heavy . the chai is ready""".split() # Count which word follows each word next_counts = defaultdict(Counter) for current, nxt in zip(corpus, corpus[1:]): next_counts[current][nxt] += 1 def next_word_probs(word): counts = next_counts[word] total = sum(counts.values()) return {w: c / total for w, c in counts.items()} print(next_word_probs("the")) # {'chai': 0.6, 'coffee': 0.2, 'rain': 0.2} print(next_word_probs("is")) # {'hot': 0.4, 'sweet': 0.2, 'heavy': 0.2, 'ready': 0.2}
And here's how it can generate text, by repeatedly picking a next word and adding it on:
def generate(start, n_words=6, seed=None): rng = random.Random(seed) words = [start] for _ in range(n_words): probs = next_word_probs(words[-1]) if not probs: # no known continuation, so stop break choices, weights = zip(*probs.items()) words.append(rng.choices(choices, weights=weights)[0]) return " ".join(words) print(generate("the")) # e.g. "the chai is sweet . the rain" (different every run)
Congratulations, that's a real language model. A terrible one, but real. It even shows you its biggest weakness. Run it a few times and sooner or later you'll get "the rain is sweet" or "the coffee is ready". When it picks the word after "is", it only looks at "is", and it has already forgotten whether the sentence was about chai, coffee, or rain.
A modern LLM fixes this in two ways:
- It looks back much further. Instead of one previous word, it considers thousands, hundreds of thousands, or in some models more than a million previous tokens. This span is called the context window.
- It uses a deep neural network instead of a counting table. A counting table can only handle word combinations it has literally seen before. A neural network generalises, so it can make sensible predictions for sentences nobody has ever written.
Everything else is scale.
§ 07 · Self-SupervisionTraining Data That Labels Itself
Supervised learning needs labelled examples, and labels are expensive. Ramesh needed a full week of his workers' time to label 2,000 mangoes, and labelling millions of leaf photos would need plant pathologists. So how could anyone label enough data to train a model with billions of parameters?
They don't have to, and this is the clever part of language modelling: the text labels itself. The "correct answer" at any position in a sentence is just the word that actually comes next, and it's already sitting right there.
Take one ordinary sentence: "Priya ordered masala chai at the station." That single sentence automatically turns into six training examples:
← swipe to see the whole figure →
Seven words give you six free training examples. This is called self-supervised learning, because the supervision signal comes from the data itself. (In practice the model learns from all of these positions at once, in a single pass. That's part of why the Transformer architecture in Article #02 is so efficient to train.)
Now scale that up. Books, websites, encyclopaedias, news articles, scientific papers, forums, and billions of lines of code all become training data, every sentence of them. Modern LLMs are trained on trillions of tokens. A large share of everything people have written and published becomes one gigantic, free, pre-labelled dataset.
What "just predicting the next word" forces a model to learn
The training objective sounds almost trivially simple. But think about what a model has to absorb to get good at it across all that text:
| What it learns | Fill in the blank | Answer |
|---|---|---|
| Grammar | "She ___ to the market yesterday" | "went", not "go" |
| Facts | "The capital of Karnataka is ___" | "Bengaluru" |
| Reasoning patterns | "If the train leaves at 6 am and the journey takes 3 hours, it arrives at ___" | "9 am" |
| Style and structure | How code is formatted, how a poem is laid out, how a formal letter opens and closes | n/a |
| Everyday common sense | "She dropped the glass on the tiled floor and it ___" | "shattered" |
Imagine a student whose only job is to fill in the next blank of every sentence in an entire library. To get really good at it, they'd have no choice but to learn the history, science, geography, and literature in those books. The task is simple. Doing it well isn't.
§ 08 · GenerationGenerating Text, One Token at a Time
A model that predicts the next token can do something remarkable: it can write. The recipe is the same one our tiny chai model used:
- Give the model some text (your prompt).
- It predicts a probability for every possible next token.
- Pick one token and add it to the end of the text.
- Feed the extended text back in and predict the next token.
- Repeat until the model produces a special "end" token or hits a length limit.
This is called autoregressive generation. "Auto" means self here, because each step builds on the model's own previous outputs. It's also why chat apps show the answer appearing word by word. That isn't a typing animation added for effect; the model really does produce its answer one token at a time.
All of this happens at inference time, the moment a trained model is being used to produce a response. (The opposite is training time, when the model is still learning.) If the loop reminds you of the predictive-text suggestions on your phone keyboard, you've got the right picture, just at a vastly bigger scale. Your keyboard suggests one likely next word based on the last word or two you typed. An LLM does the same basic thing, except it weighs your entire conversation so far, draws on patterns learned from a huge amount of human writing, and commits to one token at a time until a complete answer has formed. It's thousands of autocomplete steps in a row.
And because the model generates new text rather than just labelling existing text, LLMs are a form of Generative AI. That word deserves a proper definition, which §09 gives it.
Choosing the next token: greedy vs sampling
At step 3, how do we pick? The obvious approach is to always take the single most likely token, which is called greedy decoding. It sounds sensible, but in practice it tends to produce flat, repetitive text, and it can get stuck in loops, repeating the same phrase over and over.
The more common approach is sampling: choose randomly, but in proportion to the probabilities. Think of a lottery where each token holds tickets in proportion to its probability. "Sowing" holds 41 tickets out of every 100, "planting" holds 22, and "banana" holds a vanishingly small fraction of one. The likely words still win most of the time, but not every time, which is why clicking "regenerate" in a chatbot usually gets you a different answer.
Temperature: the creativity dial
Under the hood, a network doesn't produce probabilities directly. It produces raw scores called logits, one per token, and a function called softmax converts them into probabilities (it turns any list of scores into positive numbers that add up to 1). Before that conversion, the scores can be divided by a setting called the temperature:
- Low temperature (e.g. 0.2) exaggerates the gap between high and low scores. The top token dominates, and the output becomes focused and predictable.
- High temperature (e.g. 1.5) shrinks the gap. Less likely tokens get a real chance, and the output becomes more varied and adventurous. Push it too far and it turns incoherent.
Here's the effect on three candidate tokens:
import math def softmax_with_temperature(scores, T): exps = {w: math.exp(s / T) for w, s in scores.items()} total = sum(exps.values()) return {w: round(v / total, 3) for w, v in exps.items()} scores = {"hot": 2.0, "sweet": 1.0, "ready": 0.5} print(softmax_with_temperature(scores, T=0.5)) # {'hot': 0.844, 'sweet': 0.114, 'ready': 0.042} ← focused print(softmax_with_temperature(scores, T=1.0)) # {'hot': 0.629, 'sweet': 0.231, 'ready': 0.14} ← balanced print(softmax_with_temperature(scores, T=2.0)) # {'hot': 0.481, 'sweet': 0.292, 'ready': 0.227} ← adventurous
The scores and the model stay the same. Only the temperature changes, and "hot" drops from an 84% favourite to under 50%.
← swipe to see the whole figure →
A practical rule of thumb: use a temperature near zero when you want consistency, for example when extracting data, classifying text, or writing code. Use moderate values for conversation and higher values for brainstorming or creative writing. You'll also run into two related settings. Top-k only ever chooses among the k most likely tokens, and top-p only chooses among the smallest set of tokens whose probabilities add up to, say, 90%. Both stop the model from occasionally picking a bizarre low-probability token. You'll set these yourself when calling LLM APIs in Article #07, though not every model exposes every setting.
§ 09 · The Core IdeaWhat "Generative" Actually Means
We've used the word generative several times now, and it's the central idea of this whole series, so let's pin it down.
Judging vs creating
Look back at every model we built in §03 to §05. Ramesh's mango sorter takes two sensor readings and outputs ripe or unripe. A spam filter takes an email and outputs spam or not spam. A price model takes a Pune flat's area, floor, and locality and outputs a number. The tomato-leaf network takes a photo and outputs a disease label.
These are all discriminative models. The name comes from discriminate in its original sense, to tell things apart. They learn a mapping from inputs to outputs by drawing a boundary, fitting a curve, or learning a decision surface. Their job is to look at something and judge it.
A generative model works differently. Instead of learning a mapping from input to label, it learns what the data itself looks like. Statisticians call this the distribution of the data: which combinations are common, which are rare, and which never happen at all. A generative model learns to answer the question "What does data in this domain look like?" Once it knows that, it can sample new instances from the distribution, instances that never existed before but still feel entirely plausible.
Back to Ramesh's mangoes. The discriminative sorter only knows which side of a line a mango falls on. It knows where the border between the two groups runs, and nothing about what a typical mango looks like. A generative model of mangoes would learn the shape of the clusters themselves: that ripe mangoes tend to score around 3 for firmness and 8 for colour, say, and that a mango scoring 10 on both is essentially unheard of. It could then invent sensor readings for a realistic new mango that never rolled down the belt.
← swipe to see the whole figure →
Email works the same way. A spam classifier reads an email and tells you whether it's spam. A generative model trained on email could write you a new email from scratch, one that reads just like a real one, because the model has absorbed how real emails look and sound.
In short, a discriminative model judges data, and a generative model creates it.
You've already met one
This article has already shown you a generative model. The language model in §06 to §08 learned the distribution of human text: which token tends to follow which context, and how likely each one is. The generation loop in §08 literally draws new samples from that distribution, one token at a time. Even our tiny chai model is generative. It can produce "the coffee is ready", a sentence that appears nowhere in its training text.
The same recipe (learn the distribution, then sample from it) explains why a generative model trained on text can write essays and code. It's also why one trained on images can produce portraits of people who don't exist, and why one trained on music can compose a piece in the style of Bach.
A discriminative model carves the space of possible data into categories. A generative model learns to live inside that space.
It models the whole range of plausible outputs, not only the borders between them.
An old idea at a new scale
Generative modelling has been around for a long time. If you've studied ML before, you may have met Gaussian Mixture Models (GMMs), which describe a dataset as a blend of several overlapping bell-curve-shaped clusters, much like the density clouds in Figure 12. You may also have met Variational Autoencoders (VAEs), neural networks that compress data down to its essential features and then learn to reconstruct it, and generate new examples, from that compressed form.
What's changed in the last few years is the same thing that changed for deep learning in §05: scale, meaning more data, bigger architectures, and more compute. Push those dials far enough and the results start to feel qualitatively different. The models stop producing merely plausible data points and start producing plausible essays, programs, images, and conversations.
§ 10 · Decoding the AcronymWhat "GPT" Stands For
You now know enough to decode the most famous acronym in AI.
creates new text by sampling from what it learned (§08, §09)
trained on general text first; more training comes after (§12)
the 2017 architecture built around attention
G is for Generative. The model generates new text one token at a time, as in §08, by sampling from the distribution it learned (§09).
P is for Pre-trained. The model is first trained on a huge, general body of text before it's specialised for any particular job. Why "pre"? Because, as §12 explains, more training comes afterwards.
T is for Transformer. This is the specific neural network architecture that Google researchers introduced in 2017, in a paper titled "Attention Is All You Need." The Transformer is built around attention: when it processes each word, the model decides how much to focus on every other word in the context.
Take the sentence "The fishermen moved their boats because the bank of the river was flooding." To read "bank" correctly, the model needs to focus on "river" and "flooding" and ignore the money sense of the word. Or take "Priya lent Ananya her notes because she had missed the lecture." Who missed the lecture? The model has to connect "she" with the right person, and the rest of the sentence ("lent … her notes because …") is the clue.
Think of attention as a highlighter. As you read each word, you mentally highlight the other words in the passage that matter most for understanding it. Transformers do this for every word, in parallel, across many layers. Just as important, Transformers process every position in a sentence at the same time during training. GPUs are very good at exactly that kind of parallel work, and it's what made training these models at huge scale practical. Article #02 explains the Transformer properly, and Article #05 opens up the attention mechanism in full detail.
§ 11 · The FamilyThe Wider World of Generative AI
LLMs are the focus of this series, but they belong to a larger family. Generative AI covers several technologies, each aimed at a different modality, meaning a type of data such as text, images, audio, or video. Here's a quick tour.
← swipe to see the whole figure →
Text generation: Large Language Models
This is the category you hear about most, and the one this series focuses on. Models like GPT (you now know what all three letters mean), Claude, Gemini, and Llama are trained on huge amounts of text and can produce coherent, relevant language on almost any topic. They write, summarise, translate, reason, and hold conversations. When people say "AI" today, this is usually what they mean.
Image generation: diffusion models
Models like Stable Diffusion, Midjourney, and Google's Imagen generate images from text prompts. The most widely used technique is called diffusion. The model learns to gradually remove noise from random static until a structured image appears, with your prompt guiding each step. (Not every image generator works this way, but diffusion is the one to understand first.)
Picture a sculptor who starts with a block of pure static instead of a block of marble, like the snow on an old TV with no signal. This sculptor has studied thousands of real images closely enough to know, at every step, which specks of noise to smooth away to reveal a face, a bird, or a mountain range. Repeat that refining step a few dozen times and a coherent image emerges from what began as randomness.
← swipe to see the whole figure →
Notice that this is the §09 idea again: the model has learned the distribution of real images, and each new picture is a fresh sample from it. We won't go deep on image models in this series, but the same intuition carries over.
Audio and music
Text-to-speech (TTS) models, which read written text aloud, now produce natural-sounding voices that are hard to tell apart from real recordings. On the music side, MusicGen (from Meta) and Suno can generate full songs, with lyrics, instruments, and style, from a short text description.
Code generation
If you're heading toward engineering, this one matters to you directly. Tools like GitHub Copilot and Cursor can complete functions, generate boilerplate, explain unfamiliar code, write tests, and suggest fixes. They run on LLMs trained heavily on code, so everything you learn about LLMs in this series applies here too.
Video generation
Video is one of the fastest-moving areas. Google's Veo, Runway, and Kling generate short clips from text prompts. The hard part is temporal consistency: keeping a person's face, clothes, and surroundings stable from one frame to the next so the clip flows naturally instead of flickering. Video is harder than images because every frame has to agree with every other frame. It also illustrates the point about names: OpenAI's Sora, one of the best-known early video models, was shut down in 2026.
Multimodal models
Modern flagship models aren't limited to one modality. A single model can take in, and often produce, text, images, and sometimes audio. You can show one a photo of a whiteboard and ask it to explain the diagram, or show it a chart and ask for a written report.
Multimodality used to arrive as a separate add-on. When GPT-4 launched in 2023, for example, its ability to read images was rolled out months later as a separate "vision" version. Today it's expected of any frontier model and built in from the start, and the lines between modalities are fading.
The engine underneath
Most of the modern systems in this tour have one thing in common: the Transformer architecture and its attention mechanism, the "highlighter" from §10. Versions of the Transformer power LLMs, many image generators, and multimodal models alike. That's why Article #02 starts there.
§ 12 · Training StagesFrom Text Predictor to Assistant: The Three Training Stages
If an LLM is "just" trained to continue text, how does it end up answering questions politely, following instructions, and turning down harmful requests? Modern assistants are trained in stages, and each stage builds on the one before.
← swipe to see the whole figure →
Stage 1: Pre-training, or learning the world from text
This is everything we've covered so far: next-token prediction over trillions of tokens. It uses by far the most data and compute, typically weeks to months of training on thousands of GPUs. The result is called a base model, also known as a foundation model (§13 explains that term, because it changes how AI products get built). A base model has absorbed grammar, facts, writing styles, coding patterns, and a lot of general knowledge.
But a base model is a document continuer, not an assistant. It doesn't know it's supposed to help you. All it knows is how to continue whatever text it's given in the most plausible way. Give a base model a question and watch what happens:
PROMPT How do I reset my Wi-Fi router? A BASE MODEL MIGHT CONTINUE WITH How do I reset my Wi-Fi router? How do I change my Wi-Fi password? How do I find my router's IP address? Related questions: Why is my internet slow at night?
It isn't being unhelpful on purpose. On the internet, a question is very often followed by more questions: FAQ pages, forum listings, "people also ask" boxes. The model is faithfully continuing the kind of document it thinks it's looking at. If your prompt happened to look like the start of a help article, it might give you a real answer instead. That unpredictability is the problem.
We say the base model isn't yet aligned, meaning its behaviour doesn't reliably match what people actually want from it: to be helpful, honest, and harmless.
The good news, learned in practice, is that base models are very steerable. The knowledge is already inside them, so a fairly small amount of extra training can reshape their behaviour a lot. You're teaching format and manners, and the facts are already there.
Stage 2: Supervised fine-tuning, or learning to follow instructions
In this stage, called supervised fine-tuning (SFT) or instruction tuning, the training objective doesn't change at all: predict the next token. What changes is the data. Instead of raw internet text, the model trains on carefully curated examples of instructions paired with high-quality responses, written or reviewed by people:
- "Explain photosynthesis to a ten-year-old." → a clear, age-appropriate explanation
- "Translate this into Hindi: …" → an accurate translation
- "Summarise this email in two sentences." → a crisp summary
- Multi-turn conversations showing how to ask clarifying questions and handle follow-ups
This dataset is tiny next to the pre-training data, somewhere between thousands and perhaps millions of examples rather than trillions of tokens, because good human-written responses are expensive to produce. Here, quality matters far more than quantity.
The model picks up a new habit: when I see an instruction, respond the way a helpful assistant would. It also learns how a conversation is laid out, using special marker tokens that separate the user's message from the assistant's reply.
Stage 3: Learning from human preferences
Supervised fine-tuning has a limitation. Many requests don't have a single correct answer, and writing the perfect response is hard. Try writing the ideal message to politely decline a close friend's wedding invitation. It's surprisingly difficult. But if someone shows you two drafts, picking the better one is easy.
That observation is the basis of Reinforcement Learning from Human Feedback (RLHF):
- The model generates several different responses to the same prompt.
- Human raters compare them and mark which is better: more helpful, more accurate, safer, or better in tone.
- These comparisons are used to train a separate reward model, which learns to predict which responses people would prefer.
- The LLM is then trained with reinforcement learning (learning by trial and reward, a bit like training a dog with treats) to produce responses the reward model scores highly. It's also held close to its previous behaviour, so it doesn't learn to exploit quirks in the reward model instead of actually getting better. That failure mode is called reward hacking.
RLHF is a big part of why modern assistants are helpful, why they turn down harmful requests, and why their tone feels natural.
It isn't the only approach any more. Direct Preference Optimization (DPO) learns from the same kind of "A is better than B" comparisons but skips the separate reward model, which makes it simpler to run, and it's widely used for open-weight models. Constitutional AI, developed by Anthropic, uses a written set of principles and AI-generated feedback to scale the process up. More recently, reinforcement learning on tasks with checkable answers, such as maths problems with known solutions or code that has to pass tests, has been used to train reasoning models, which work through problems step by step before answering (more on those in §16).
The new-employee analogy
Here's one way to keep all three stages in your head at once:
- Pre-training is a brilliant graduate who has read an entire library. They know an enormous amount, but they have no idea how your company works or what you expect from them.
- Supervised fine-tuning is onboarding. They study worked examples of good emails, reports, and customer replies until they can produce work in the expected format.
- Preference tuning is their first year of feedback from a manager: "This reply was better than that one, and here's why." Over time their judgement and tone sharpen.
§ 13 · The BridgeFoundation Models: How GenAI Changes the Way AI Gets Built
This section connects the classic ML of §03 to §05 with everything else in this series.
The base model from Stage 1 has another name you'll hear constantly: foundation model. A foundation model is a large, general-purpose model trained once on a huge, broad dataset and then adapted, rather than retrained from scratch, for many different jobs. The name was chosen on purpose. Like the foundation of a building, one model is meant to hold up many different structures built on top of it (chatbots, coding assistants, translators, classifiers) without anyone redoing the foundational work each time.
| Traditional ML | Foundation models (GenAI) | |
|---|---|---|
| Scope | Task-specific | General-purpose |
| Training data | Labelled datasets | Massive unlabelled data (self-supervised, §07) |
| Output | A label or a number | Text, images, code, audio |
| Adaptation | Collect new data and train a new model | Prompt it, or fine-tune it |
| Model count | One model per task | One model, many tasks |
| Example | Ramesh's mango sorter | One LLM that summarises, translates, drafts emails, and writes code |
Two ways of building
The traditional ML workflow goes: define a task → collect labelled data → train a model → deploy → repeat for the next task. Ramesh needed a week of hand-labelling for one sorting job. If he now wants to spot bruised mangoes as well, he needs a fresh labelled dataset and a new model. Every task gets its own model, its own dataset, and its own training run, which is expensive and slow to scale.
Foundation models turn this around. One model is pre-trained on huge amounts of unlabelled data and learns rich, general representations of language (or images, or code). To use it for a specific task, you have two options:
- Prompt it. Describe the task in plain language, perhaps with a few examples. §16 shows how.
- Fine-tune it. For more specialised performance, keep training it briefly on a small, focused dataset so it gets noticeably better at one kind of task. Stage 2 in §12 did this for general helpfulness; here you aim it at your job. Think of a general-practice doctor doing a specialist residency. The medical degree stays, and a focused layer of expertise goes on top.
← swipe to see the whole figure →
The big change is going from "train a model per task" to "prompt one model for many tasks."
That changes the economics of building AI products. Work that used to take months of data collection, labelling, training, and deployment can often be prototyped in hours, and the cost of trying out an AI-powered feature is far lower than it used to be.
Don't dismiss traditional ML
Everything you learned in §03 to §05 still matters, and you'll keep using it. When you have clean, structured data in rows and columns (customer churn, fraud detection, credit scoring), traditional ML models are often faster, cheaper, and easier to explain than an LLM, and they need less data. The usual tools here are logistic regression, random forests, and XGBoost, a fast, widely used library for gradient-boosted decision trees.
Ramesh's sorter is a perfect example. It has two sensor readings, a clear label, and a decision to make within milliseconds on a factory floor. It also has to run at the edge, meaning directly on a small device beside the conveyor belt instead of making a round trip to a data centre. A tiny trained classifier on a cheap chip beats a 70-billion-parameter language model on speed, cost, and reliability. When you have well-defined labels and a narrow, stable task, the traditional ML toolkit is the right one.
§ 14 · Check YourselfTest Your Mental Model
You now know enough to predict why LLMs can do many of the things they do. Before reading each answer below, stop and try to answer it yourself using the ideas above.
Why can an LLM summarise a long document?
The training data is full of texts followed by their summaries: news stories with headlines, research papers with abstracts, long forum posts ending in "TL;DR" (internet shorthand for "too long; didn't read", followed by a one-line summary). So the pattern long text → short version was learned during pre-training. When you ask for a summary, the whole document sits in the model's context window, so it can pay attention to all of it while generating. Supervised fine-tuning then sharpened the skill with deliberate examples.
Why can it translate between languages?
Nobody built it as a translator. But the web holds vast amounts of parallel text: bilingual websites, the same Wikipedia article in dozens of languages, subtitles, official documents published in several languages. Predicting the next word across all of that produces translation ability as a side effect.
Why can it write code?
Pre-training data includes huge amounts of public code and programming discussion. Code is also highly structured and consistent, so in many ways it's easier to predict than casual conversation.
Why can it answer general-knowledge questions?
Two stages work together. The knowledge came from pre-training, and the habit of answering questions helpfully (instead of continuing with more questions) came from fine-tuning and preference training.
Why does it sometimes stumble when counting the letters in a word, or on precise arithmetic?
It sees tokens, not individual letters (Article #03 explains why "strawberry" causes trouble), and it produces plausible text rather than working through a calculation digit by digit. The engineering fix is to let the model use tools such as a calculator or a code interpreter, which we build in Article #24.
§ 15 · HallucinationsWhen the Model Makes Things Up
You've probably seen it happen. An LLM confidently states a "fact" that's wrong, invents a book that doesn't exist, cites a research paper nobody wrote, or suggests a software function that isn't real. This is called hallucination: fluent, confident output that is false.
Why it happens
Everything you've learned in this article helps explain it:
- The training objective rewards plausibility, not truth. During pre-training the model is scored on whether its predicted token matches the text. Nothing checks whether a statement is true. If a plausible-sounding completion exists, the model is naturally inclined to produce it.
- The training data sounds confident. Encyclopaedia entries, textbooks, and news articles rarely say "I'm not sure, but…". The model learns an authoritative voice and uses it whether or not it actually knows the answer.
- Knowledge is compressed, not stored. An LLM has no database of facts to look things up in. Its knowledge is spread as patterns across billions of parameters. Common facts come back reliably, while rare ones (the population of a small town, the authors of an obscure paper) come back fuzzy. It's like recalling a phone number you saw once: you remember its shape, and you might confidently produce ten digits with several of them wrong.
- Knowledge has a cut-off date. A model only knows what was in its training data, which was collected up to a certain date, its knowledge cut-off. It simply doesn't know about anything that happened later, and it may answer from outdated information without warning you that it's stale.
Newer models generally hallucinate less than older ones, and fine-tuning can teach a model to say "I don't know" more often. Progress isn't perfectly steady, though. OpenAI's own testing in 2025 found that its o3 reasoning model made things up more often on one factual benchmark than the older o1 had. No current model is immune, so treat hallucination as an engineering constraint to design around rather than a bug that will go away if you wait.
The fix engineers actually use: grounding
What makes grounding work: an LLM is far more reliable at using information you place in front of it than at recalling information from memory. Anything in the prompt sits right there in the context window, where attention can find it directly. Anything learned in pre-training has to be reconstructed from compressed patterns.
It's the difference between a closed-book exam and an open-book one. When accuracy matters, give the model the book.
Suppose a user asks "What is the current RBI repo rate?" The Reserve Bank of India changes its policy rate from time to time, so a model answering from memory might give a figure that was right at its training cut-off and wrong today. Instead, we can build a prompt like this:
Use only the information in the context below to answer the question. If the answer is not in the context, reply "I don't know." Context: [paste the latest RBI monetary policy press release here] Question: What is the current RBI repo rate?
Now the model's job has changed from remembering to reading comprehension, which it's very good at. This is called grounding: anchoring the model's answer in source material you supply instead of letting it generate freely from memory.
← swipe to see the whole figure →
Search-enabled assistants work this way. They run a web search, paste the most relevant results into the model's context, and then generate an answer, often with citations back to the sources. Apply the same pattern to your own documents, such as company policies, product manuals, or research papers, and it's called Retrieval-Augmented Generation (RAG). RAG is one of the most important patterns in GenAI engineering. We build it properly in Phase 5 of this series, using the embeddings from Article #04 and the vector databases from Phase 3.
§ 16 · Emergence and PromptingSurprising Abilities, and How to Use Them
Scale and emergence
As researchers trained ever-larger models on ever more data, something unexpected happened. Models started showing abilities nobody had explicitly trained them for, like translating between language pairs that rarely appear side by side, doing basic arithmetic, following instructions for tasks they had never seen, and writing convincingly in a requested style.
Some researchers noticed that certain abilities seemed to switch on suddenly once a model passed a certain size, and called them emergent abilities. Others have argued that some of that suddenness comes from how we measure. An all-or-nothing score (the answer is either exactly right or it isn't) can make steady, gradual improvement look like an abrupt jump. Either way, the broader lesson holds: bigger models trained on more data tend to be more capable, often in ways nobody specifically aimed for.
Three of these abilities matter a great deal in everyday use, and each one follows directly from thinking of an LLM as a text continuer trained on human writing.
Zero-shot: just describe the task
You can often ask a model to do something completely new with nothing but an instruction. This is called zero-shot prompting, because you provide zero examples. Say a colleague sends you this:
hey sir cant come tmrw, fever, will send the report by eve
and you ask the model: "Rewrite this as a polite, formal leave-request email to my manager." It'll produce a proper email with a subject line, greeting, explanation, and sign-off, even though it was never trained on this exact task. It combines patterns it learned separately: formal email conventions, what a leave request contains, and how to expand informal abbreviations.
The main tip for zero-shot prompting is to be specific. State the audience, the tone, the length, and the format you want.
Few-shot: show, don't just tell
Sometimes instructions alone aren't enough, especially when you need a precise output format. The fix is what you'd do for a new colleague: show them some examples. This is few-shot prompting. Say you need to clean up messy Indian phone numbers into one consistent format:
Convert each phone number to the format +91-XXXXX-XXXXX. Input: 9876543210 Output: +91-98765-43210 Input: 098765 43210 Output: +91-98765-43210 Input: +91 91234 56789 Output: +91-91234-56789 Input: 91234-56780 Output: ▌ the model continues: +91-91234-56780
The model will almost always continue with +91-91234-56780.
Why does this work so well? Because at heart an LLM is still a text continuer. You've built a document with a clear, repeated pattern, and the most probable continuation of that document is to follow the pattern one more time. Notice that the last "Output:" is left blank on purpose. That makes the natural next text the answer itself, in your format, with nothing extra.
This ability to pick up a task from examples inside the prompt, without any change to the model's parameters, is called in-context learning.
Chain-of-thought: give the model room to think
Try this question: "Priya buys 3 notebooks at ₹45 each and pays with a ₹500 note. She spends half of her change on pens. How much money does she have left?"
If you ask for just the final number, a model may jump straight to a guess and sometimes get it wrong. Now add a simple instruction: "Work through this step by step, then give the final answer." You'll typically get something like:
Cost of notebooks: 3 × ₹45 = ₹135 Change received: ₹500 − ₹135 = ₹365 Spent on pens: ₹365 ÷ 2 = ₹182.50 Money left: ₹365 − ₹182.50 = ₹182.50
The answer is far more likely to be correct. This technique is called chain-of-thought prompting, and it works for two reasons, both of which follow from what you now know:
- Each token gets a roughly fixed amount of computation, one pass through the network. Squeezing a multi-step problem into a single answer token asks too much of that one pass. Spreading the reasoning over many tokens spreads the computation over many passes.
- Everything the model writes becomes part of its context. Once "₹365" is written down, the next step can simply attend to it (§10) instead of having to hold it implicitly. The written steps act as working memory, the model's version of scratch paper.
← swipe to see the whole figure →
Modern reasoning models take this much further. They're trained with the reinforcement learning on checkable answers mentioned in §12, and they write out a long internal chain of reasoning before giving their final answer. That makes them much better at maths, coding, and logic problems. The trade-off is that all that reasoning is made of tokens, and tokens cost time and money (Article #03 shows how to measure that).
§ 17 · Use CasesGenAI in the Real World
You now know how these models work, where they fall short, and how to steer them. So where are they actually being used? Here are six areas, each with a before-and-after example.
Developer tools
Other uses: code review feedback, documentation, explaining legacy code, suggesting bug fixes.
Content and writing
Other uses: summarising articles, marketing copy, drafting emails, translation that keeps the original tone.
Customer support
Other uses: FAQ assistants, routing tickets to the right team, flagging angry customers who need a human quickly.
Data and analysis
Other uses: summarising reports, pulling structured fields out of messy documents such as invoices and contracts, explaining unusual spikes in data.
Search and knowledge (RAG)
It's the grounding pattern from §15 at company scale: an open-book exam where the relevant pages are found automatically and handed to the model before it answers. We go deep on RAG in Phase 5 of this series. It's one of the most useful patterns in GenAI engineering.
Creative work and games
Two patterns run through all six examples. First, a person stays in the loop. The change is rarely "AI replaces the person"; it's usually "the person reviews a draft instead of starting from a blank page". Second, the structure is always the same:
Context goes in, relevant generation comes out.
Most of the craft of building GenAI applications is in engineering that context carefully, and that's what most of this series is about.
§ 18 · An Open QuestionIs It "Just" Predicting the Next Word?
Before we look ahead, here's an open question that serious researchers still disagree about.
In one well-known 2023 study, a model trained only on sequences of moves from the board game Othello, and never shown a picture of the board, turned out to be tracking the state of the board internally, because doing so helped it predict legal next moves.
Both views may be partly right. "Next-token prediction" describes the training objective. It doesn't necessarily limit what a model has to learn in order to meet that objective. Saying a chess grandmaster is "just choosing the next move" is technically accurate, and it tells you nothing about how much understanding it takes to choose well.
§ 19 · RoadmapSeries Roadmap: What You're Building Toward
Here's the full map, so you know where we're going and why each step matters. The series is organised into seven phases, and each phase builds a complete layer of the stack before the next one starts.
Exact article numbers may shift a little as the series grows, because new topics sometimes earn an article of their own. The seven-phase structure is fixed, and this roadmap will stay up to date as we go.
Where this article's ideas go next
Almost every idea you met today gets its own deep dive later:
| Idea from this article | Where it's covered in depth |
|---|---|
| The Transformer and attention (§10, §11) | Article #02, then Article #05 |
| Tokens and vocabulary (§06) | Article #03 |
| Embeddings (§04) | Article #04, then Phase 3 |
| Zero-shot, few-shot, chain-of-thought (§16) | Article #06 |
| Temperature and sampling settings (§08) | Article #07 |
| Measuring whether a model is actually good (§15, §18) | Article #10 |
| Fine-tuning and preference training (§12, §13) | Articles #11 and #12 |
| Bias and harmful outputs (§07) | Article #13 |
| Letting the model use tools like a calculator (§14) | Article #24 |
| Grounding and RAG (§15, §17) | Phase 3 and Phase 5 |
The end goal: by the time you finish this series, you'll be able to design, build, and deploy a full Generative AI application from scratch, and you'll understand every layer of the stack, not just the API calls.
Prerequisites: none for this article. From Article #02 onward some comfort with Python helps, because the code examples are in Python. The ML you need to follow along is in §03 to §05. If you'd like to go deeper on classic ML first, the ML Fundamentals Series is the place to start.
§ 20 · SummaryThe Mental Model You Now Have
Seventeen things to take with you
- AI ⊃ ML ⊃ Deep Learning ⊃ LLMs. AI is the goal. Machine learning is the approach of learning from examples, deep learning is ML with multi-layer neural networks, and LLMs are deep learning models trained on text.
- Machine learning learns rules from data instead of having them hand-written. Features go in, labels come out, and training minimises the loss.
- Generalisation is the real goal. A model that memorises its training data (overfitting) fails on new data, so always evaluate on a held-out test set.
- Classic ML hit a wall on unstructured data. Images and text have huge numbers of inputs with very complex relationships, and hand-crafted feature engineering couldn't keep up.
- Neural networks stack layers of simple neurons with non-linear activation functions in between. Depth lets them learn their own features automatically, from simple to complex.
- Training uses gradient descent (walking downhill on the loss) and backpropagation (assigning blame backwards through the network) to tune millions or billions of parameters.
- A language model predicts the next token. That's a classification problem with as many classes as there are tokens in the vocabulary.
- Self-supervised learning lets any text serve as training data, because the next word is its own label. Modern LLMs are trained on trillions of tokens.
- Generation is autoregressive: predict a token, append it, repeat. Sampling and temperature control how predictable or varied the output is.
- Discriminative models judge, generative models create. A discriminative model learns the boundary between categories. A generative model learns the distribution of the data itself and can sample brand-new, plausible examples from it.
- GPT stands for Generative Pre-trained Transformer. The Transformer's attention mechanism lets every word focus on the other words that matter most.
- Generative AI covers many modalities: text (LLMs), images (mostly diffusion), audio and music, code, video, and multimodal models that combine them. Transformers are the engine underneath most of them.
- Assistants are built in stages: pre-training (knowledge) → supervised fine-tuning (instruction following) → preference tuning such as RLHF or DPO (helpfulness, safety, judgement).
- Foundation models change how AI gets built, from "train a model per task" to "prompt (or fine-tune) one model for many tasks". Traditional ML is still the right tool for structured data, narrow tasks, and fast on-device predictions.
- Hallucinations happen because models are trained for plausibility rather than truth, hold compressed rather than stored knowledge, and have a knowledge cut-off. Grounding, supplying relevant source material in the prompt, is the main fix and the basis of RAG.
- Zero-shot, few-shot, and chain-of-thought prompting all work because an LLM is a text continuer that treats everything in its context as usable information.
- Real-world GenAI follows one pattern: context goes in, relevant generation comes out, and a person usually reviews the result. Engineering that context well is what this series teaches.
Key terms at a glance
| Term | Plain-English meaning |
|---|---|
| Feature | A measurable input the model uses (firmness, a pixel value, a word embedding) |
| Label | The correct answer for a training example |
| Classification / Regression | Predicting a category / predicting a number |
| Loss | A single number measuring how wrong the model's prediction was |
| Overfitting | Memorising the training data instead of learning a pattern that generalises |
| Neural network | Layers of simple weighted-sum units with non-linear "bends" between them |
| Parameter | One adjustable number (a weight or a bias) inside a model |
| Gradient descent | Repeatedly nudging parameters "downhill" to reduce the loss |
| Backpropagation | Working backwards from the error to compute how each parameter should change |
| Token | A chunk of text (often a word or part of a word) that the model reads and writes |
| Context window | The maximum amount of text (in tokens) the model can consider at once |
| Self-supervised learning | Training where the labels come from the data itself, e.g. the next word |
| Autoregressive | Generating one token at a time, each based on everything produced so far |
| Temperature | A setting that makes output more focused (low) or more varied (high) |
| Inference | Using a trained model to produce an output, as opposed to training it |
| Discriminative model | A model that judges or labels its input (spam / not spam, ripe / unripe) |
| Generative model | A model that learns what the data looks like and can create new examples of it |
| Modality | A type of data: text, images, audio, video |
| Diffusion model | A generator, most often for images, that turns random noise into a picture step by step |
| Multimodal model | One model that handles several modalities, such as text and images together |
| Base / Foundation model | A large pre-trained model that continues text but doesn't yet behave like an assistant; the general-purpose starting point many applications are built on |
| Fine-tuning | Briefly training an existing model further on a small, focused dataset |
| SFT / Instruction tuning | Fine-tuning on curated instruction–response pairs |
| RLHF | Reinforcement Learning from Human Feedback: training on human preference comparisons |
| Alignment | Making a model's behaviour match what people intend: helpful, honest, harmless |
| Hallucination | Fluent, confident output that is false |
| Knowledge cut-off | The date after which the model has no training data |
| Grounding / RAG | Supplying source material in the prompt so answers are based on it |
| Zero-shot / Few-shot | Prompting with instructions only / with a few worked examples |
| Chain-of-thought | Prompting the model to reason step by step before answering |
📝 Learner exercise
Use any chatbot you have access to and run three small experiments, then finish with a quick sorting task. After each one, write a single sentence explaining the result with a concept from this article.
- Sampling. Ask the same creative prompt three times ("Write a two-line poem about monsoon traffic"), then ask the same factual question three times ("What is the boiling point of water at sea level?"). Which one varies more between attempts, and why?
- Hallucination and grounding. Ask about something obscure that you personally know well, like your school's founding year or a local landmark's history. Check the answer. Then paste a reliable source into the prompt, ask again, and compare the two answers.
- Chain-of-thought. Write a three-step word problem of your own. Ask for "just the final answer" in one chat and "work through it step by step" in a fresh chat, then compare the results.
- Judge or create? Label each of these as discriminative or generative, and name the modality: your email's spam folder · your bank's "was this you?" fraud alert · your phone's face unlock · the next-word suggestions on your phone keyboard · an app that turns a text prompt into a poster image. (Hint: one of these is sneakier than it looks. Re-read the "neat twist" in §09.)
The Transformer Architecture
You now have the whole story, from sorting mangoes to chatting with an assistant, plus a map of where the series goes from here. From now on, we open up each piece of the machinery you met today, starting with the most important one.
We'll see why the attention mechanism was such a big breakthrough, and why it solved problems that the older ways of reading text in order couldn't (those were recurrent neural networks, or RNNs, and their improved cousins, LSTMs). Along the way, the "highlighter" from §10 becomes a proper intuition for self-attention, and every LLM concept after that will feel grounded instead of magical.
It's the article where everything starts to click. See you there.