a question with no settled answer · and it matters more every year

Intelligence is the ability to solve a problem you have never seen before

passes, but surely isn’t

fails, but surely is

What isintelligence?

Every definition anyone has offered lets in something that obviously is not intelligent, and shuts out something that obviously is. That is not a failure of effort — four centuries of people have tried. This page walks the definitions, breaks each one, and then goes somewhere more uncomfortable: the place where AI sits, and where nobody has a good answer.

Start with a crow

A crow is shown a narrow glass tube. At the bottom, out of reach, is a small bucket of food with a handle. Beside the tube lies a straight piece of wire.

The crow picks up the wire, wedges one end, bends it into a hook, lowers it into the tube, catches the handle, and lifts the food out. It had not been taught this. Step through what actually had to happen.

what the crow had to do

Almost nobody hesitates about this one. Whatever intelligence is, that looked like it.

So: say what it was. Not what it felt like — what property the crow demonstrated that a rock, a calculator or a thermostat does not have. Every attempt below is a real answer that serious people have given. Each one works until it doesn’t.

question one · answer before you read on

A different crow watches the first one, then does the same thing on its first try. Which crow showed more intelligence?

There is a defensible case for each. Pick the one you actually believe before revealing.

the answer, such as it is There is no settled answer, and the disagreement is the interesting part. Invention from scratch looks more impressive, but social learning is rarer in nature and arguably harder — it needs a model of what another creature is doing and why. Meanwhile the third option is the one a benchmark would pick, because outcomes are the only thing a test can score, and it is obviously missing something.

Hold on to that gap between what we admire and what we can measure. It is the crack that everything else on this page falls through.

Four definitions, and what breaks them

Take each seriously. Each has been the mainstream answer at some point, and each is still defended today. Pick one and try to hold it.

hold a definition and see what it lets in

the case for it

what it wrongly admits

Notice the shape of the failure. It is never that a definition is simply wrong. It is that every definition has two failure modes at once: something passes that should not, and something fails that should not. Tighten it to exclude the slime mould and you exclude the crow. Loosen it to include the crow and the thermostat walks in.

A test with both a false positive and a false negative is not a definition. It is a rough guide with a marketing problem.

question two

A slime mould — a single-celled organism with no brain and no neurons — grows through a maze and finds the shortest route to food. Is that intelligence?

This is a real experiment, repeated many times. Commit to an answer.

the answer, such as it is Most biologists would say the second. Most people who then have to say why discover they cannot, without also ruling out things they want to keep.

“It is just chemistry following a gradient” is true — and a neuron is also chemistry following a gradient, in vastly greater numbers. If the reply is “yes, but there are far more of them, arranged far more intricately”, then you have quietly changed your definition from a kind of thing to an amount of a thing, and you now owe an answer to: how much is enough?

The third option is not a dodge. It may be the honest one, and by the end of this page you may prefer it.

What the tests actually measure

If defining it is hard, perhaps measuring it is easier. This is what the twentieth century tried.

The observation that started it: people who do well on one kind of mental task tend to do well on others. Vocabulary, spatial puzzles, arithmetic, memory — the scores correlate. That correlation was named g, and much of psychometrics has been an argument about what it is.

The uncomfortable part is not that the tests are useless. They predict some real outcomes, at population scale, better than chance. The uncomfortable part is what happens when you ask what they are made of.

what a test can and cannot reach
A test is a proxy, and every proxy can be optimised against. That is not a flaw in these particular tests. It is what proxies are. Keep it in mind for two sections’ time, when the thing being tested is a machine that is very good at being optimised.
question three

Someone trains for six months specifically on IQ test questions and raises their score by fifteen points. Did they become more intelligent?

the answer, such as it is Almost everyone says the second, immediately and confidently. Which is worth pausing on, because it means we all already believe the score is not the thing — we believe there is a real underlying quantity that the test is only pointing at.

That belief is doing enormous work, and it is not obviously justified. If the test is not the thing, what is the thing, and how would you check?

Now apply exactly this reasoning to a model that was trained on material resembling the benchmark it is later scored on. If you said “no” above, consistency demands you say “no” there too.

The goalposts have always moved

There is a pattern in this field old enough to have a name. Every time a machine does something that was said to require intelligence, that thing stops counting as intelligence.

Step through it and watch the definition retreat.

things that required intelligence, until they didn’t

Two readings of that pattern, and they are not compatible.

The cynical reading: we are moving the goalposts to protect our specialness. Each time, we redefine intelligence as whatever machines cannot do yet, which makes the claim unfalsifiable and slightly pathetic.

The generous reading: we keep learning something real. Every time a task falls, we discover it did not need the thing we thought it needed — and that is genuine knowledge. Chess turned out not to require judgement. That is a discovery about chess, not a retreat.

Both readings are partly right, which is why the argument never ends. The useful question is not “are we moving the goalposts” but “what did we learn about the task when it fell?”
question four

If a machine eventually does every single thing on that list and everything you can add to it, would that settle the question?

the answer, such as it is Your choice here reveals which kind of definition you have been carrying without stating it.

Choose the first and you hold a behavioural definition: intelligence is what a thing does, full stop. Clean, testable, and it commits you to calling any sufficiently capable system intelligent regardless of what is inside.

Choose the second or third and you hold a constitutive one: intelligence is about what is going on inside, and behaviour is only evidence for it. Also defensible — and it owes an account of what the inside has to be like, which nobody has ever managed to give without smuggling in “like ours”.

Neither is naive. But you cannot hold both, and most arguments about AI are two people holding one each without noticing.

The things that turned out to be hard

Here is the result that most damaged the intuition. Rank a set of tasks by how hard they feel to a person. Then rank them by how hard they turned out to be for a machine. The two orderings are close to inverted.

Drag the line and compare.

hard for a person, or hard for a machine?

Roboticists noticed this decades ago: the things we find effortless — walking over uneven ground, recognising a face in bad light, picking up a mug — are computationally enormous, while the things we find effortful and prestigious, like formal logic and chess, are comparatively cheap.

The usual explanation is evolutionary. Perception and movement have been under refinement for hundreds of millions of years, and are so heavily optimised that they feel like nothing from the inside. Abstract reasoning is recent, thin, and expensive — which is exactly why it feels like work, and why we mistook the feeling of effort for the presence of intelligence.

We built our definition out of the things that feel hard to us. There is no reason the universe should have organised difficulty around human introspection, and it did not.

Where AI actually sits

Now the part everyone actually came for, and where the piece stops being comfortable.

A modern language model does things that, on any of the four definitions above, count. It solves problems it was not shown. It uses language fluently. It adapts its output to a goal you state. It transfers, at least somewhat, across domains it was not explicitly trained on.

It also does things that make people certain it does not count. It states falsehoods with complete confidence. It can fail at arithmetic a child can do. Its competence is oddly shaped — brilliant here, absent one step away — in a pattern no person exhibits.

Both descriptions are accurate. Test a claim about it and see which way the evidence actually points.

a claim, and what the evidence does to it

The honest summary is that the evidence does not resolve. It is not that we lack data — we have enormous quantities. It is that the data cannot settle a question whose terms were never fixed.

question five · the uncomfortable one

Suppose you become convinced a system genuinely understands. What observation could have convinced you — and what observation would change your mind back?

Both halves matter. A belief with no possible disconfirming evidence is not a belief about the world.

the answer, such as it is Most people find they are in the second or third group, including people with strong public views in both directions.

That is worth sitting with, because it means the disagreement is often not empirical at all. Two people can agree on every fact about what a system does and still disagree completely, because they are disagreeing about what the word should mean — and no experiment adjudicates a definition.

If you are in the first group, you have something most of the debate lacks. Write it down before you read on; it is more valuable than an opinion.

Three questions people keep merging

A great deal of confusion comes from treating these as one question. They are separate, they have different answers, and mixing them makes conversations unwinnable.

Place a few things on the grid.

capable, and conscious, are different axes
more capable ↑
more likely conscious →

The grid is uncomfortable on purpose. Most people place an insect low-left and a language model high-left — capable, and not conscious. But almost nobody can justify the horizontal position of either without appealing to how similar the thing is to us, which is not a principle so much as a preference.

Capability is measurable. Experience is not observable from outside, in anything, including other people.

We extend the assumption of inner experience to other humans on the basis of similarity, not evidence. That assumption is almost certainly correct. It is also, strictly, an inference we cannot check — and it is the only tool we have for anything else.

The grey area, stated plainly

Here is what is genuinely unresolved, without the usual confidence in either direction.

Is intelligence one thing or many?

The correlation between mental abilities is real, which supports one general factor. But the pattern of AI capability is powerful evidence for the other view: we now have systems that are superhuman at some cognitive tasks and useless at others, with no correlation whatever between them. If a single general capacity were required, that should be impossible. It is not merely possible — it is what every system we have built looks like.

Does it require a body?

One serious position holds that intelligence is inseparable from acting in a physical world, and that a system trained only on descriptions has learned the shadow of understanding rather than the thing. Another holds this is a failure of imagination about what statistics over enough description can reach. Both are held by careful people. Neither has produced a test the other accepts.

Does the substrate matter?

If a system were built from something other than biology and behaved indistinguishably from a person across every situation, would it be intelligent? Say yes and you accept that intelligence is a pattern, indifferent to what it runs on. Say no and you owe an account of what carbon supplies that silicon cannot — and every attempt so far reduces to “it is not made of the stuff we are made of”.

question six

Which of these would most change your view if it turned out to be true?

the answer, such as it is There is no right choice. The point of the question is that you can name one — which puts you ahead of most of this debate, where positions are usually held without any stated condition for revision.

Whichever you picked, notice that it is an empirical claim. It could be checked. That makes it a far better thing to argue about than the word itself, which cannot be.

What survives

Nothing on this page produced a definition. That is the honest result, not a failure to try hard enough, and four centuries of effort suggest it is not about to yield.

But something useful survives, and it is more practical than a definition would have been. Replace the question you cannot answer with three you can.

the question you cannot answer, and the ones you can

ask this instead

because

None of those three requires knowing what intelligence is. All three are answerable with evidence. And all three are what you actually needed when you asked the original question — whether you were deciding what to deploy, what to trust, or what to worry about.

The word was never the point. “Is it intelligent?” feels like a question about the system. It is usually a question about what we are willing to grant, and that is a decision, not a discovery.

One last thing, and it is the reason this matters beyond philosophy. We are now building systems whose capabilities we cannot fully characterise, and deploying them into decisions that affect people, using a word we cannot define to describe what they are doing. The definitional problem was tolerable when it was only about crows.

the last question · no answer is provided

Having read all of that: has your own definition changed, or only your confidence in it?

This one has no reveal. It is the only question on the page whose answer belongs to you.

pick one — nothing is hidden behind this