AI, machine learning, neural networks and deep learning really do sit inside one another, like India, Maharashtra, Mumbai, Bandra. Then the diagram keeps going — and everything after that point is a different kind of word entirely. Keep pressing the button and watch where it stops working.
People keep asking some version of: is a chatbot AI, or is it machine learning, or is it GenAI? The honest answer is unsatisfying and completely correct: yes, all of them, at the same time.
That is not a dodge. It is the same kind of answer as asking whether someone lives in India, or Maharashtra, or Mumbai, or Bandra. Nobody finds that confusing. Climb the ladder and watch what changes.
All four are true at once. They are not four different places you might live; they are four levels of zoom on one place. You would use different ones in different conversations — “India” to someone abroad, “Bandra” to someone across town — and switching between them is not a contradiction.
These four are settled and agreed on. Notice that the dates go down as the rings get smaller — each inner ring is a later, more specific answer to the problem the outer ring set. Pick one.
The example that proves the outer ring is bigger than learning: in 1997 a chess machine beat the world champion. Unquestionably AI, and a landmark. It contained no machine learning whatsoever — hand-written evaluation rules plus enormous brute-force search. Nothing about it improved on its own.
This is the correction that matters most, and it is the opposite of what the press implies. The techniques in the second ring — regressions, decision trees, random forests, nearest neighbours — were almost all invented between the 1950s and the 1990s. None is a neural network. And they still run most of the world’s actual production machine learning.
There is a solid practical reason. Pick the shape of your data.
Anyone proposing a neural network for a churn model should be asked, politely, what they tried first.
The third ring — neural networks — is a way of learning that stacks layers of very simple units. The honest analogy is an assembly line of extremely dim workers. No individual worker understands the product. Each does one trivial thing and passes it on. The complexity comes entirely from how many there are and how they are wired, not from any one being clever.
The history is worth walking, because it corrects a false impression of inevitability. Step through it.
So what changed? Not the idea. The idea was already there, and had been for half a century. Data and computing power arrived. That is a genuinely useful lesson for judging any technology today: the same idea can be worthless and then transformative without changing at all.
Deep learning is neural networks with a large number of layers, which turns out to change what is possible, not merely how well it works. And here is the single most important technical idea on this page:
Classical machine learning needs a human to decide what to look at. Deep learning works out what to look at, by itself.
Same task both times: is this photo a dog? Switch between the two approaches.
That human step in the classical column has a name — feature engineering — and for thirty years it was the actual job. Deep learning largely automated it away.
Up to deep learning, every word is the same kind of word: a category of technique, each a genuine subset of the one before. After deep learning, the words are different kinds of things entirely, and stacking them in one picture makes them look like a hierarchy when they are not.
Cricket fixes this instantly. Walk down the chain and watch where it stops being a nesting.
Is a particular team “inside” a format? Loosely, in a way you would accept in conversation. But it is not the same kind of thing at all, and drawing it as one more nested circle implies something false — that every match in that format involves that team, which is nonsense.
The famous diagram does exactly this. It takes an architecture, a capability, a class of model, a brand, a version and a product, and draws all six as though they were nested categories.
One distinction here costs organisations real money, so it deserves a minute. A trained model is an engine. A chat product is a car. Switch between buying the two.
Three consequences follow. Behaviour differs — the same underlying model answers differently in a product versus through the raw interface, because the product wraps it in instructions you never see. You are usually buying a product, not a model, so comparing one chat product against another mixes categories. And what your company builds sits at the model layer, which is where most of the actual value and most of the actual risk lives.
The fastest way to see that the inner circles are not really nested is to hold up things that do not fit. Each of these is real, well known, and snaps one specific arrow.
The image generator and the text classifier are the pair to sit with. Together they prove the arrows point in both directions: there is generative AI that is not built on the transformer design, and there are transformers that generate nothing at all. Two things that overlap but where neither contains the other cannot be drawn as one circle inside another.
Beginners assume the rings are trivia. They are not — they are a cost and risk map. Which ring your problem sits in determines almost everything about how you would resource it. Pick a ring and price it.
Three things worth saying out loud from that. The foundation-model column has no build cost and no end to its running cost — a genuinely different economic shape from anything most organisations have bought before, and budgets built for one-off software projects handle it badly.
The explainability row is why banking and insurance still live in the second ring. That is not conservatism. If you must justify a declined loan to a regulator, an unreadable model is not an option.
Being in the outer, older ring is not a defeat. A simple model that ships this quarter and can be explained beats a deep learning project that lands next year and cannot.
Twelve jobs. Each is described by what it does, never by the technique behind it — because you cannot place something on a diagram if the description itself is jargon you would have to decode first. How far in does each one go?
The thermostat, the chess computer, the image generator and the ticket sorter are where the real learning is. The thermostat is on none of these rings — it is plain automation, and someone always wants to put it inside AI. The chess computer proves the outer ring is bigger than learning. And the last two, side by side, show why generative AI cannot be drawn as a circle inside anything: one generates and is not a language model; the other is built on the same underlying design as a chatbot and writes nothing at all.
Two honest limits, because a diagram this tidy invites overconfidence.
It shows techniques, not capability. New inner circles have appeared every few years and will keep appearing. Press the button and watch the innermost ring move.
Nothing in this diagram tells you how far the field will go, and any slide claiming otherwise is selling something. What the history does tell you is narrower and more useful: the newest circle is always the one that looks inevitable in hindsight and looked like a dead end for decades beforehand.
The leaf names drift constantly. Particular model families and version numbers change on a timescale of months. Treat any specific product name as illustrative and check a live source before quoting current capabilities or rankings. The shape is stable; the labels at the edges are not.
One more thing worth heading off. “Should we move from machine learning to AI?” is a sentence that gets said in real meetings, and it does not mean anything — nobody moves from ML to AI, because they were already doing AI. Usually the sentence means “we should use large language models”, which is a specific and expensive choice that deserves to be argued on its merits rather than smuggled in through vocabulary.