Is Saying an LLM Doesn't Think Like Saying a Calculator Can't Do Numbers?
Where the calculator analogy for LLMs holds and where it breaks: what interpretability, chain-of-thought and philosophy of mind say about thinking.
Update, October 2026. I clarified one term through this note; the body is unchanged. The reason is in the update section near the end.
Why this matters
I build and evaluate systems based on language models, and the question “but do these models actually think?” runs through every serious conversation on the subject: the client deciding how much autonomy to give an agent, the colleague dismissing everything as autocomplete, and the paper announcing emergent capabilities. The answer shapes concrete choices: how much to delegate, how to verify outputs, and what vocabulary to use in technical documentation. A quip that has circulated for some time compresses the dispute into a single sentence. Taking it apart, piece by piece, is the most honest way I know to arrive at a defensible position.
The calculator as analogy, and as trap
“Saying an LLM doesn't think is like saying a calculator can't do numbers.” The line circulates as a joke, and like good jokes, it contains a compressed argument. It is worth unpacking carefully because the analogy is both more robust and more fragile than it first appears. Understanding where it holds and where it breaks says something important about language models and about the people judging them.
Start with the calculator. The sentence “a calculator doesn't know numbers” is, in a precise sense, true. A calculator has no concept of number: it switches electrical states according to rules fixed by its designer, and nothing in its operation resembles the understanding a child acquires while learning to count. Yet the same sentence, said to someone using a calculator to file their taxes, sounds absurd. The calculator does arithmetic: it performs many routine arithmetic operations reliably, often faster and more accurately than a human. The absurdity comes from using “knowing” in two different senses: one constitutive, meaning possessing understanding, and one functional, meaning correctly performing a function. The sentence is true in the first sense and false in the second.
Someone claiming that “an LLM doesn't think” often makes the same shift without stating it. They begin with a defensible constitutive premise — there is no evidence of consciousness or intentionality in the strong sense of the term — and use it to suggest a far broader functional conclusion: that there is no inference, no abstraction, and nothing that deserves the vocabulary of reasoning. The calculator quip makes this move visible. But making it visible is not the same as refuting it. At this stage, the controversy does not arise from a technological difference. It arises from a linguistic ambiguity. The rest of this article asks whether a substantive question remains once that ambiguity is removed. My answer is yes — but not the one either faction expects.
What an LLM actually does
To argue honestly, I first need to establish what is known. On this point, the situation has changed considerably since 2020: some mechanisms in some models and tasks are now partially understood, although no frontier model is comprehensively explained.
The starting point is familiar: an LLM is trained to predict the next token over enormous text corpora. From this correct but incomplete description comes the slogan “it's just autocomplete.” The slogan omits what training produces: high-dimensional distributed representations in which concepts are encoded in superposed, compositional forms. These representations support non-trivial in-context generalization. Mechanistic interpretability identified, as early as 2022 in the work of Olsson and colleagues on induction heads, candidate circuits — attentional mechanisms that copy and complete patterns — whose causal contribution to in-context learning has been tested through interventions in specific model settings.
The picture is more nuanced for explicit reasoning, and it requires careful wording. Wei and colleagues showed in 2022 that prompting a model to produce intermediate steps (chain-of-thought) markedly improves performance on arithmetic, commonsense, and symbolic tasks. Kojima and colleagues showed that the bare instruction “let's think step by step,” without examples, can elicit latent capabilities. These results are solid as behavioral phenomena. They do not show that the verbalized steps are the internal mechanism through which the model reaches its answer. Here, the skeptic has a strong empirical argument: a growing literature on chain-of-thought unfaithfulness, opened by the work of Turpin and colleagues, documents cases in which the verbalized explanation does not reflect the computation that produced the answer. A model can reach a result by one route and narrate another. Add the fragility of reasoning under superficial perturbations: irrelevant reformulations of a problem can degrade performance that, if it rested on robust abstract competence, should remain unaffected. Anyone dismissing chain-of-thought as linguistic theater therefore has data to cite, not only intuitions.
The reply is not to deny these data, but to put them in context. The work of Dutta and colleagues on the mechanistic interpretability of chain-of-thought shows that models deploy multiple parallel neural pathways for step-by-step reasoning. These results provide evidence of structured internal computation in the models and tasks studied, even when the verbal account is unfaithful. Unfaithful self-reports are also familiar in cognitive psychology: humans readily confabulate post-hoc rationalizations. This does not prove that LLMs reason as humans do. It shows that “the verbal explanation is unfaithful” does not mean “there is no underlying computation worth calling reasoning.”
Finally, there is residual opacity. The work of Elhage and colleagues on superposition explains why interpretability is difficult: models compress more concepts than they have available dimensions, producing polysemantic neurons that resist direct interpretation. Opacity is neither magic nor mystery. It is an architectural property with known, partly tractable causes. The honest balance is this: some mechanisms in some models and tasks are now partially understood, although no frontier model is comprehensively explained; whether these mechanisms constitute a form of genuine reasoning is plausible but contested; whether they constitute thought is a question that data alone cannot settle. To understand why, I need to change terrain.
What "thinking" means
The terrain is philosophical, and the first thing I find there is that the verb “to think” has no neutral definition that the parties could agree on before examining the data. Every position carries its own criteria for attribution and therefore reaches different verdicts on the same facts.
Functionalism, in the formulation Putnam gave it in 1967, holds that mental states are defined by the causal role they occupy — by their relations to inputs, outputs, and other states — rather than by the substrate that realizes them. If pain is what pain does, then it can be realized in a brain, in a circuit, and in principle in any system with the appropriate causal organization. For a consistent functionalist, the question about LLMs is empirical: do they have the functional organization of thought, or not? The silicon substrate is not, by itself, an argument.
Searle's Chinese Room (1980) is the classic challenge to this framework. A person manipulates Chinese symbols by following rules, without understanding Chinese, and produces answers indistinguishable from those of a speaker. Searle therefore concludes that syntactic manipulation is not sufficient for understanding, and that no program, qua program, can understand anything. This is a constitutive objection, not a behavioral one: it challenges the idea that function exhausts the mental. The replies are well known. The strongest, the systems reply, argues that it is not the person in the room who must understand, but the overall system of which that person is a component. After forty-five years, the debate remains open. That is the relevant fact: if a thought experiment from 1980 still divides philosophers, the concept of understanding is not settled enough to serve as an arbitrating criterion.
Chalmers introduced a useful distinction into the recent debate: the question of phenomenal consciousness — whether there is something it is like to be the system — is separable from the question of cognitive capabilities. His conclusion about current LLMs is probabilistic and cautious: they are most likely not conscious, though the possibility cannot be ruled out for future systems. The distinction matters here in a negative sense. Anyone denying thought to LLMs by appealing to the absence of consciousness must first argue that thought requires consciousness. That is a respectable thesis, but it is far from obvious given how much human thought occurs without conscious accompaniment.
Then there is Wittgenstein, who addresses the question “can a machine think?” directly in the Philosophical Investigations (§§359–360) and defuses it: “only of a human being and what resembles (behaves like) a living human being can one say: it thinks”. His point is not that thinking machines are metaphysically impossible. It is that “thinking” belongs to a grammar: a network of uses, criteria, and forms of life. Extending the term to new entities, or withholding it from them, is not a discovery about the world but a decision about language. Millière and Buckner, in their two-part survey (part one and part two), show how contemporary debates about grounding, compositionality, and world models in LLMs revisit philosophical controversies that were never resolved. Technical novelty did not bring with it the criteria needed to judge it.
The lesson is not skepticism but structure. Anyone saying that “an LLM doesn't think” — or that it “thinks” — is applying a theory of thought, even when they believe they are stating a fact. The available theories diverge precisely at the points needed to decide the case.
Where the analogy breaks
Here I need to break the opening analogy deliberately, because its limit is more instructive than its strength.
When you say that a calculator “does arithmetic without knowing it,” you can do so with complete precision because arithmetic is a fully formalized domain. It has an exact syntax of operations, an exact semantics of results, and public, shared criteria of correctness. It is known what the calculator realizes, how it realizes it, and that the mechanism exhausts the task: nothing remains outside what is called “computing correctly.” For that reason, you can isolate what it lacks — understanding — and observe that the function does not require it. The statement about the calculator is comparatively precise because arithmetic has formal rules and public criteria of correctness.
Nothing equivalent exists for thought. There is no shared definition, no formalization, and no verification criterion independent of competing theories. The previous section showed why: functionalists, Searleans, and Wittgensteinians do not diverge on the data. They diverge over what would count as thought. The statement “an LLM doesn't think” therefore cannot have the same epistemic status as “a calculator doesn't understand numbers.” The second is an observation within a formalized domain; the first claims comparable precision in a domain that does not possess it. To know that something does not think, you would need to know what thinking is. There is no consensus definition that settles the question for models, and the concept remains contested even in accounts of human thought.
But this argument cuts both ways. If the lack of a theory of thought makes “doesn't think” undecidable, it also makes “thinks” undecidable. Anyone using the disanalogy only against skeptics while ignoring it before enthusiasts would be practicing strategic agnosticism, which is a vice rather than a position. The consequence is symmetric: neither attribution is, under current conditions, a scientific finding. Both are proposals — more or less motivated and more or less useful — about how to extend a concept whose conditions of application were never fixed for cases like this. The middle position recently defended by Tayyar Madabushi, Torgbi and Bonial, which describes LLM capabilities as context-directed extrapolation over training data priors — beyond the stochastic parrot, short of human reasoning — is valuable precisely because it rejects the false dichotomy. Yet its vocabulary remains a descriptive choice, not a verdict on essence.
The decisive difference, then, is not between calculators and LLMs. It is between a formalized domain, where function and understanding can be separated precisely, and a non-formalized domain, where every attribution carries an undeclared theory.
The real gaps
Nothing I have said licenses triumphalism. This section is here to prevent it. Current LLMs have deep limits that should be described as architectural properties, rather than brandished as slogans in either direction.
A base language model has no persistence across contexts; any apparent memory or durable goal comes from surrounding system components. As I explain in every conversation starts from zero, what the interface presents as memory is context re-inserted from outside, not sedimented experience. It has no agency of its own: it forms no goals that survive the context window, and it has no history of interactions with the world that constrains its future dispositions. It has no body: its relation to reality is mediated entirely by text. This restates, in technical terms, the older grounding problem, forcefully posed by Bender and Koller: whether the symbols it manipulates ever touch anything other than more text. Shanahan proposed describing the conversational behavior of these systems as role-play: the model is not an interlocutor with a stable point of view, but a generator of plausible characters. Treating it as a unitary subject is a category error encouraged by the conversational interface. This proposal is useful as an antidote to naive anthropomorphism, even if, taken literally, it risks reducing functionally real computation to theater.
These limits are real and documented. What remains unproven is whether they are constitutive of thought or merely features of its human variant. Persistence, embodiment, and biographical continuity: are these properties of thought as such, or properties of the only available exemplar of a thinker? Anyone treating them as definitive arguments assumes the latter answer while presenting it as the former. Again, this is theory presented as observation. The honest formulation is this: the current limits of LLMs accurately describe their architecture; whether they also establish the impossibility of thought depends on a definition of thought that nobody possesses. It is also possible — and remains an open question — that architectural developments will erode some of these limits, returning the question under different conditions.
Wittgenstein had moved the question
I can now return to the opening quip and weigh it properly. Yes, saying that an LLM doesn't think resembles saying that a calculator can't do numbers: in both cases, a true statement about understanding is used to imply a false conclusion about function. But the resemblance ends where formalization ends. The question can be closed for the calculator; it cannot be closed for LLMs. Not because data are missing, but because the definition the data would need to satisfy or violate is missing.
This is why the controversy is not waiting for a decisive experiment. No measurement, benchmark, or interpretability result can settle the broader philosophical question without an agreed operational definition of ‘thinking.’ The dispute is not about facts alone. It is about the grammar assigned to a verb calibrated on human beings and now applied to cases for which it was not designed. Wittgenstein saw this clearly: asking whether a machine can think is not simply proposing a hypothesis to verify; it is negotiating a use. If the answer changes one day — and it might — it will not be only because models have changed. It will also be because the working meaning of “thinking” has changed, as it did when “memory” extended to computers and “intelligence” to tests.
In the meantime, both shortcuts remain shortcuts. “It's just a statistical parrot” denies functions that are observable and partly explained. “It thinks like a human” attributes what no criterion licenses. The defensible position is the uncomfortable middle: these are systems that realize functions for which, in other contexts, the vocabulary of thought would be used. They do so on a radically different substrate, with real gaps, within a conceptual debate that machines have reopened and that only a change in how the term is used could resolve, while empirical work can still narrow the practical questions.
Transparency note: this article was co-produced with a large language model, through a structured workflow of research, drafting and supervised revision. Given the thesis of the piece, I invite the reader to consider this circumstance materially relevant — one way or the other.
Update — October 2026
The section on where the analogy breaks calls the calculator claim decidable and the LLM claim undecidable. I did not mean the computability-theory sense of those words. I meant that the calculator claim can be checked against shared, public criteria of correctness, and that no such criteria exist yet for thought. The key takeaway now uses that wording.
References
- Bender & Koller (2020), Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data, ACL 2020.
- Bender, Gebru, McMillan-Major & Shmitchell (2021), On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?, FAccT 2021.
- Wei et al. (2022), Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.
- Kojima et al. (2022), Large Language Models are Zero-Shot Reasoners.
- Olsson et al. (2022), In-context Learning and Induction Heads, Transformer Circuits Thread.
- Elhage et al. (2022), Toy Models of Superposition, Transformer Circuits Thread.
- Turpin et al. (2023), Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, NeurIPS 2023.
- Dutta, Singh, Chakrabarti & Chakraborty (2024), How to Think Step-by-Step: A Mechanistic Understanding of Chain-of-Thought Reasoning.
- Mirzadeh et al. (2024), GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.
- Shanahan, McDonell & Reynolds (2023), Role-Play with Large Language Models.
- Chalmers (2023), Could a Large Language Model be Conscious?.
- Millière & Buckner (2024), A Philosophical Introduction to Language Models: Part I and Part II.
- Searle (1980), Minds, Brains, and Programs; SEP entry: The Chinese Room Argument.
- Putnam (1967), Psychological Predicates; SEP entry: Functionalism.
- Wittgenstein (1953), Philosophical Investigations, §§359–360; SEP entry: Ludwig Wittgenstein.
- Tayyar Madabushi, Torgbi & Bonial (2025), Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors.
FAQ
Is an LLM just autocomplete?
The description “it predicts the next token” is correct but incomplete. Training produces distributed representations in which concepts are encoded in superposed, compositional forms, and mechanistic interpretability has identified candidate circuits, such as induction heads, whose causal contribution to in-context learning has been tested through interventions in specific model settings. The slogan omits precisely this part.
Does chain-of-thought prove that an LLM reasons?
Not by itself. Intermediate steps improve performance as a behavioral phenomenon, but the unfaithfulness literature shows that a verbalized explanation may not reflect the computation that produced the answer. At the same time, these results provide evidence of structured internal computation in the models and tasks studied: an unfaithful account does not imply the absence of underlying computation.
Why is the calculator a different case from an LLM?
Arithmetic is a fully formalized domain with public criteria of correctness, so function and understanding can be separated precisely. For thought, there is no shared definition and no verification criterion independent of competing theories. As a result, neither “thinks” nor “doesn't think” has the status of a scientific finding.
What limits do current LLMs have?
A base language model has no persistence across contexts; any apparent memory or durable goal comes from surrounding system components. It has no agency of its own, and it has no body, because its relation to reality is mediated by text alone. These are real, documented architectural properties. Whether they are constitutive of thought, rather than properties of its human variant, remains unproven.
Related articles
mklang: The Document Is the Program
A declarative .mkl file for LLM-driven state machines. Models generate. The machine decides what happens next. What the language guarantees and does not.
Oct 2, 20267 min read#mklang#LLM#DSL#State-Machine#AgentsWhy Autonomous Research Agents Hallucinate — and How a Critic Loop Surfaces Unsupported Claims
Research agents can look authoritative and still invent citations. A critic with source access catches unsupported claims — it doesn't erase risk.
Nov 15, 202410 min read#AI Agents#Research#LLM#ProductionEmotional Memory
I check this article against the public proof at DOI 10.5281/zenodo.19972258.
Oct 4, 20261 min read#Memory#LLM#Research#State