# Proof of Work vs Inference: Cost as Signal, Not Identity

> Inference is not proof of work. I separate two costly, checkable spends of compute: a rule for scarce consensus, and a model plus prompt for utility.

Published: 2026-10-02
Canonical: https://gianlucamazza.it/en/blog/pow-vs-inference
Tags: Inference, Cost, Eval, LLM, Production

## Why this matters

Inference is not proof of work.

I keep hearing the two names in the same sentence, as if a costly forward pass were a new consensus trick. The rhyme is real. The identity is not. Both spends burn compute that someone else can check. They do it for opposite reasons, against different objects, and they fail in different ways. If I flatten them, I rename a utility bill as a protocol and lose track of what the spend was for.

This is a teaching cut, not a product claim. I measure gates, record rejects, and keep retrieval receipts. I do not treat those habits as a reinvented proof of work.

Two different spends of compute:

- **Proof of work** is costly, checkable work against a **rule**. The work is meant to be useless. The scarcity is the point: consensus needs a ticket that is hard to mint and cheap to verify.
- **Inference** is costly, checkable work against a **model and a prompt**, plus an eval or a gate when I am not guessing. The spend aims at utility: tokens and watts toward an answer, not a lottery ticket.

The shared axis is cost as a signal. The forbidden public claim is that the two are the same thing.

## What proof of work actually is

Proof of work is a filter on who may append to a shared log. A participant spends energy on a search that has no use beyond satisfying a public rule: find a nonce such that the hash of the block header falls under a difficulty target. Anyone can re-run the check in milliseconds. Almost no one can mint a fresh valid ticket without paying the search cost again.

Three properties matter, and they are narrower than "this looks expensive, so it must be serious":

1. **The check is against a rule.** The verifier does not ask whether the work was insightful. The verifier asks whether the output meets a predetermined predicate.
2. **The work is deliberately useless.** If the search produced a valuable side effect, an attacker could harvest that effect without caring about consensus. The waste is how the protocol stops the ticket from being a by-product of ordinary useful work.
3. **The cost buys scarce consensus, not vague credibility.** A valid proof admits a block, or it does not. It does not award a reputation score or a feeling that the operator "paid their dues."

People borrow "proof of work" when they mean "I spent money, therefore trust me." That is not the protocol. A costly gesture can be a signal and still fail every consensus property: no shared rule, no cheap public check, no agreement about what the ticket admits.

The check is also one-sided in a way inference is not. Once the rule is met, the work is done. There is no second question about whether the nonce was a good nonce. Goodness is the predicate.

## What inference actually is

Inference is a forward pass. I load a model, I supply a prompt, I spend memory bandwidth and energy, and I receive tokens. If I am operating the system rather than watching a demo, I then check the output against a schema, a frozen eval set, a policy gate, or a human review. The spend happens before that second check, which is what makes it useful or wasted.

The object of the check is not a puzzle. It is a model plus a prompt, and then whatever gate I put after the decoder:

- **Against the model and the prompt:** did this runtime, with these weights and this context, produce this sequence? That part is checkable. I already write about that habit in [bank-grade eval harnesses](/en/blog/bank-grade-agent-evals).
- **Against an eval or a gate:** did the sequence satisfy the job? That part can fail after the tokens have already been ordered.

I do not run inference to mint a scarce ticket. I run it because I want an answer, a tool call, a patch, a retrieval-grounded paragraph. The cost is real. The aim is utility. If the eval fails, I still spent the compute. Cost was real; the answer was not.

That is why "inference is expensive, therefore it is proof of work" is a category error. Expense is not a protocol. A lottery ticket and a lab assay can cost the same and still not be the same instrument.

A local GGUF runtime and a hosted API are both inference. Neither becomes a consensus rule because I can quote a throughput number. I refuse unmeasured tok/s here on purpose. If a figure is not in a published artifact, it does not belong in this article.

## Same axis, opposite intent

The useful overlap fits in one table. I keep it visible so the metaphor cannot quietly expand.

|                        | Proof of work                                                            | Inference                                                                |
| ---------------------- | ------------------------------------------------------------------------ | ------------------------------------------------------------------------ |
| Check against          | A rule or puzzle                                                         | A model and a prompt, then an eval or gate                               |
| Intent                 | Deliberately useless work for scarce consensus                           | Costly work that aims at a useful answer                                 |
| Cost as signal         | Energy and compute buy a ticket that is hard to mint and cheap to verify | Energy and compute buy a candidate output; the gate decides if it counts |
| Forbidden public claim | —                                                                        | Inference is proof of work                                               |

Same rhyme: scarce, verifiable work. Opposite intent: waste-on-purpose versus a useful answer.

Cost as signal is the part I keep. If a step is free, I cannot tell whether anyone ran it. If a step has a bill, I can ask what the bill purchased. That question is legitimate in both pictures. It does not make an LLM stack a consensus protocol.

The failure modes split the same way. A bad proof of work is an invalid ticket, so the log does not move. A bad inference can still be a valid decode: the model produced those tokens, and the gate then says they do not do the job. I paid for a candidate, and the candidate failed.

## Useful-PoW is a concept, not a protocol

There is a bridge idea, and I want it named so it does not sneak in as a product.

**Useful-PoW** is the thought that the same checkable spend could also do a job someone wanted anyway: verifiable work that is not pure waste. It remains a concept. It is not a product I ship, and it is not a name I put on xllama, mklang, or ragfs.

Two limits stay in place:

1. **A useful side effect does not invent a consensus rule.** If I cannot say what the ticket admits, who verifies it, and what happens when two valid tickets conflict, I have a workload, not a protocol.
2. **Inference alone is not a consensus protocol.** A model plus a prompt can be costly and checkable and still say nothing about who appends to a shared log. An eval or a price tag does not close that gap.

I use the bridge to keep two pictures in view, not to weld them. The [calculator analogy](/en/blog/llm-thinking-calculator) already showed that a rhyme can be exact in one place and empty in another. Cost is the exact place. Identity is the empty one. I will not flatten this into "we reinvented proof of work."

## Practice: cost, ledger, and retrieval honesty

The practice I can stand behind is narrower than a Useful-PoW pitch. I keep three honest books in public artifacts: what the forward pass cost, what I already ordered when a gate refuses the output, and what I actually retrieved.

### xllama: measure the cost gate, do not rename it

[xllama](https://github.com/gianlucamazza/xllama) is my local inference project on Xbox developer-mode hardware: GGUF and ONNX Runtime GenAI, CPU or GPU chosen per workload, no retail store path. The [Xbox Series S report](/en/blog/slm-on-xbox) is the narrative receipt. The [generated benchmark summary](https://github.com/gianlucamazza/xllama/blob/main/docs/benchmarks.md) is the command-level receipt, each figure tied to the command that produced it. I do not repeat a tok/s number here.

What I take from that work is the gate, not a slogan. Local inference has a bill in watts, memory, and wall time. I can measure it and publish the command. I cannot turn that bill into proof of work by calling the measurement a puzzle.

Xbox is a Microsoft trademark. xllama is an independent research project, not a Microsoft product and not affiliated with Microsoft.

### mklang: a reject does not erase the tokens I ordered

[mklang](https://github.com/gianlucamazza/mklang) is [a declarative DSL for LLM-driven state machines](/en/blog/mklang). I use it on this site as a selective runtime for propose-and-repair loops; the accepted decision is [ADR-0003](https://github.com/gianlucamazza/personal-website/blob/main/docs/adr/0003-mklang-content-runtime.md). The machine can halt, repair, or escalate. [Deterministic gates still decide](/en/blog/ai-workflow-engines). The [live validation log](https://github.com/gianlucamazza/personal-website/blob/main/docs/mklang-live-validation.md) records input and output tokens on those runs, including a copy-review rewrite later marked `corrected`.

That is ledger honesty, not a charge product. Refusing bad output does not erase the tokens I already ordered. The spend happened before the verdict. The [mklang project card](/en/projects#mklang) is the public index for the repository. I am not calling mklang a Useful-PoW product.

### ragfs: retrieval has a bill; receipts beat "trust me"

[RAGFS](https://github.com/Venere-Labs/ragfs) mounts an indexed directory and exposes operations, results, and undo through the filesystem. I wrote the interface decision in [Why I Gave Agents a Filesystem Instead of Another API](/en/blog/agent-filesystem). Search on that mount runs a local embedding model and stores vectors in LanceDB. The first query pays a download and an index; later queries pay the path the [public documentation](https://venere-labs.github.io/ragfs/) describes. I have not published a retrieval-quality benchmark for this setup, so I do not invent one here.

Retrieval is not a free preface to inference. It has a bill, and it has a result I can read: which path I asked for, what the JSON said, whether an `undo_id` exists. That is a receipt. "Trust me, the context was relevant" is not. The diagnostic order I use when retrieval is weak remains [RAG in Production](/en/blog/rag-systems-production). None of that is proof of work. It is scarce, verifiable work with the opposite intent.

## Cost rhymes; equivalence doesn't

Cost rhymes. Equivalence does not.

Proof of work spends compute against a rule so that consensus can stay scarce. Inference spends compute against a model and a prompt so that someone can use the output. Useful-PoW is the concept that those two pictures might one day share a workload. It is not a name for a local runtime, a copy-review machine, or a FUSE mount.

The related posts I want next to this one are the receipts, not a scene: the [Xbox cost gate](/en/blog/slm-on-xbox), the [filesystem interface for retrieval](/en/blog/agent-filesystem), the [RAG diagnostic order](/en/blog/rag-systems-production), and the [eval harness](/en/blog/bank-grade-agent-evals). Read them as cost, ledger, and retrieval honesty.

## FAQ

### Is inference a form of proof of work?

No. Inference spends compute against a model and a prompt for a useful answer. Proof of work spends compute against a rule for scarce consensus, and the work is deliberately useless. They share a rhyme. They are not the same protocol.

### What does "cost as signal" mean here?

If a step is free, I cannot tell whether anyone ran it. If a step has a bill, I can ask what it purchased and whether a later gate accepted the result. That signal does not make inference into proof of work.

### What is Useful-PoW in this article?

A concept only: verifiable work that also does a job someone wanted. It is not a product I ship, and not a name I put on xllama, mklang, or ragfs. Inference alone is still not a consensus protocol.

### Why mention xllama, mklang, and ragfs?

They are public receipts for three honesty habits: measure the local cost gate, record tokens already ordered when a gate refuses output, and keep retrieval results instead of "trust me" context.

### Does a failed eval erase the compute I spent?

No. The forward pass happens before the gate. If the eval fails, the cost was real and the answer was not. That is ledger honesty, not a proof-of-work ticket.
