Every definition anyone has offered lets in something that obviously is not intelligent, and shuts out something that obviously is. That is not a failure of effort — four centuries of people have tried. This page walks the definitions, breaks each one, and then goes somewhere more uncomfortable: the place where AI sits, and where nobody has a good answer.
A crow is shown a narrow glass tube. At the bottom, out of reach, is a small bucket of food with a handle. Beside the tube lies a straight piece of wire.
The crow picks up the wire, wedges one end, bends it into a hook, lowers it into the tube, catches the handle, and lifts the food out. It had not been taught this. Step through what actually had to happen.
Almost nobody hesitates about this one. Whatever intelligence is, that looked like it.
So: say what it was. Not what it felt like — what property the crow demonstrated that a rock, a calculator or a thermostat does not have. Every attempt below is a real answer that serious people have given. Each one works until it doesn’t.
A different crow watches the first one, then does the same thing on its first try. Which crow showed more intelligence?
There is a defensible case for each. Pick the one you actually believe before revealing.
Take each seriously. Each has been the mainstream answer at some point, and each is still defended today. Pick one and try to hold it.
Notice the shape of the failure. It is never that a definition is simply wrong. It is that every definition has two failure modes at once: something passes that should not, and something fails that should not. Tighten it to exclude the slime mould and you exclude the crow. Loosen it to include the crow and the thermostat walks in.
A test with both a false positive and a false negative is not a definition. It is a rough guide with a marketing problem.
A slime mould — a single-celled organism with no brain and no neurons — grows through a maze and finds the shortest route to food. Is that intelligence?
This is a real experiment, repeated many times. Commit to an answer.
If defining it is hard, perhaps measuring it is easier. This is what the twentieth century tried.
The observation that started it: people who do well on one kind of mental task tend to do well on others. Vocabulary, spatial puzzles, arithmetic, memory — the scores correlate. That correlation was named g, and much of psychometrics has been an argument about what it is.
The uncomfortable part is not that the tests are useless. They predict some real outcomes, at population scale, better than chance. The uncomfortable part is what happens when you ask what they are made of.
Someone trains for six months specifically on IQ test questions and raises their score by fifteen points. Did they become more intelligent?
There is a pattern in this field old enough to have a name. Every time a machine does something that was said to require intelligence, that thing stops counting as intelligence.
Step through it and watch the definition retreat.
Two readings of that pattern, and they are not compatible.
The cynical reading: we are moving the goalposts to protect our specialness. Each time, we redefine intelligence as whatever machines cannot do yet, which makes the claim unfalsifiable and slightly pathetic.
The generous reading: we keep learning something real. Every time a task falls, we discover it did not need the thing we thought it needed — and that is genuine knowledge. Chess turned out not to require judgement. That is a discovery about chess, not a retreat.
If a machine eventually does every single thing on that list and everything you can add to it, would that settle the question?
Here is the result that most damaged the intuition. Rank a set of tasks by how hard they feel to a person. Then rank them by how hard they turned out to be for a machine. The two orderings are close to inverted.
Drag the line and compare.
Roboticists noticed this decades ago: the things we find effortless — walking over uneven ground, recognising a face in bad light, picking up a mug — are computationally enormous, while the things we find effortful and prestigious, like formal logic and chess, are comparatively cheap.
The usual explanation is evolutionary. Perception and movement have been under refinement for hundreds of millions of years, and are so heavily optimised that they feel like nothing from the inside. Abstract reasoning is recent, thin, and expensive — which is exactly why it feels like work, and why we mistook the feeling of effort for the presence of intelligence.
Now the part everyone actually came for, and where the piece stops being comfortable.
A modern language model does things that, on any of the four definitions above, count. It solves problems it was not shown. It uses language fluently. It adapts its output to a goal you state. It transfers, at least somewhat, across domains it was not explicitly trained on.
It also does things that make people certain it does not count. It states falsehoods with complete confidence. It can fail at arithmetic a child can do. Its competence is oddly shaped — brilliant here, absent one step away — in a pattern no person exhibits.
Both descriptions are accurate. Test a claim about it and see which way the evidence actually points.
The honest summary is that the evidence does not resolve. It is not that we lack data — we have enormous quantities. It is that the data cannot settle a question whose terms were never fixed.
Suppose you become convinced a system genuinely understands. What observation could have convinced you — and what observation would change your mind back?
Both halves matter. A belief with no possible disconfirming evidence is not a belief about the world.
A great deal of confusion comes from treating these as one question. They are separate, they have different answers, and mixing them makes conversations unwinnable.
Place a few things on the grid.
The grid is uncomfortable on purpose. Most people place an insect low-left and a language model high-left — capable, and not conscious. But almost nobody can justify the horizontal position of either without appealing to how similar the thing is to us, which is not a principle so much as a preference.
Capability is measurable. Experience is not observable from outside, in anything, including other people.
We extend the assumption of inner experience to other humans on the basis of similarity, not evidence. That assumption is almost certainly correct. It is also, strictly, an inference we cannot check — and it is the only tool we have for anything else.
Here is what is genuinely unresolved, without the usual confidence in either direction.
The correlation between mental abilities is real, which supports one general factor. But the pattern of AI capability is powerful evidence for the other view: we now have systems that are superhuman at some cognitive tasks and useless at others, with no correlation whatever between them. If a single general capacity were required, that should be impossible. It is not merely possible — it is what every system we have built looks like.
One serious position holds that intelligence is inseparable from acting in a physical world, and that a system trained only on descriptions has learned the shadow of understanding rather than the thing. Another holds this is a failure of imagination about what statistics over enough description can reach. Both are held by careful people. Neither has produced a test the other accepts.
If a system were built from something other than biology and behaved indistinguishably from a person across every situation, would it be intelligent? Say yes and you accept that intelligence is a pattern, indifferent to what it runs on. Say no and you owe an account of what carbon supplies that silicon cannot — and every attempt so far reduces to “it is not made of the stuff we are made of”.
Which of these would most change your view if it turned out to be true?
Nothing on this page produced a definition. That is the honest result, not a failure to try hard enough, and four centuries of effort suggest it is not about to yield.
But something useful survives, and it is more practical than a definition would have been. Replace the question you cannot answer with three you can.
None of those three requires knowing what intelligence is. All three are answerable with evidence. And all three are what you actually needed when you asked the original question — whether you were deciding what to deploy, what to trust, or what to worry about.
One last thing, and it is the reason this matters beyond philosophy. We are now building systems whose capabilities we cannot fully characterise, and deploying them into decisions that affect people, using a word we cannot define to describe what they are doing. The definitional problem was tolerable when it was only about crows.
Having read all of that: has your own definition changed, or only your confidence in it?
This one has no reveal. It is the only question on the page whose answer belongs to you.
pick one — nothing is hidden behind this