Proof of Work vs Inference: Cost as Signal, Not Identity
Inference is not proof of work. I separate two costly, checkable spends of compute: a rule for scarce consensus, and a model plus prompt for utility.
Why this matters
Inference is not proof of work.
I keep hearing the two names in the same sentence, as if a costly forward pass were a new consensus trick. The rhyme is real. The identity is not. Both spends burn compute that someone else can check. They do it for opposite reasons, against different objects, and they fail in different ways. If I flatten them, I rename a utility bill as a protocol and lose track of what the spend was for.
This is a teaching cut, not a product claim. I measure gates, record rejects, and keep retrieval receipts. I do not treat those habits as a reinvented proof of work.
Two different spends of compute:
- Proof of work is costly, checkable work against a rule. The work is meant to be useless. The scarcity is the point: consensus needs a ticket that is hard to mint and cheap to verify.
- Inference is costly, checkable work against a model and a prompt, plus an eval or a gate when I am not guessing. The spend aims at utility: tokens and watts toward an answer, not a lottery ticket.
The shared axis is cost as a signal. The forbidden public claim is that the two are the same thing.
What proof of work actually is
Proof of work is a filter on who may append to a shared log. A participant spends energy on a search that has no use beyond satisfying a public rule: find a nonce such that the hash of the block header falls under a difficulty target. Anyone can re-run the check in milliseconds. Almost no one can mint a fresh valid ticket without paying the search cost again.
Three properties matter, and they are narrower than "this looks expensive, so it must be serious":
- The check is against a rule. The verifier does not ask whether the work was insightful. The verifier asks whether the output meets a predetermined predicate.
- The work is deliberately useless. If the search produced a valuable side effect, an attacker could harvest that effect without caring about consensus. The waste is how the protocol stops the ticket from being a by-product of ordinary useful work.
- The cost buys scarce consensus, not vague credibility. A valid proof admits a block, or it does not. It does not award a reputation score or a feeling that the operator "paid their dues."
People borrow "proof of work" when they mean "I spent money, therefore trust me." That is not the protocol. A costly gesture can be a signal and still fail every consensus property: no shared rule, no cheap public check, no agreement about what the ticket admits.
The check is also one-sided in a way inference is not. Once the rule is met, the work is done. There is no second question about whether the nonce was a good nonce. Goodness is the predicate.
What inference actually is
Inference is a forward pass. I load a model, I supply a prompt, I spend memory bandwidth and energy, and I receive tokens. If I am operating the system rather than watching a demo, I then check the output against a schema, a frozen eval set, a policy gate, or a human review. The spend happens before that second check, which is what makes it useful or wasted.
The object of the check is not a puzzle. It is a model plus a prompt, and then whatever gate I put after the decoder:
- Against the model and the prompt: did this runtime, with these weights and this context, produce this sequence? That part is checkable. I already write about that habit in bank-grade eval harnesses.
- Against an eval or a gate: did the sequence satisfy the job? That part can fail after the tokens have already been ordered.
I do not run inference to mint a scarce ticket. I run it because I want an answer, a tool call, a patch, a retrieval-grounded paragraph. The cost is real. The aim is utility. If the eval fails, I still spent the compute. Cost was real; the answer was not.
That is why "inference is expensive, therefore it is proof of work" is a category error. Expense is not a protocol. A lottery ticket and a lab assay can cost the same and still not be the same instrument.
A local GGUF runtime and a hosted API are both inference. Neither becomes a consensus rule because I can quote a throughput number. I refuse unmeasured tok/s here on purpose. If a figure is not in a published artifact, it does not belong in this article.
Same axis, opposite intent
The useful overlap fits in one table. I keep it visible so the metaphor cannot quietly expand.
| | Proof of work | Inference | | ---------------------- | ------------------------------------------------------------------------ | ------------------------------------------------------------------------ | | Check against | A rule or puzzle | A model and a prompt, then an eval or gate | | Intent | Deliberately useless work for scarce consensus | Costly work that aims at a useful answer | | Cost as signal | Energy and compute buy a ticket that is hard to mint and cheap to verify | Energy and compute buy a candidate output; the gate decides if it counts | | Forbidden public claim | — | Inference is proof of work |
Same rhyme: scarce, verifiable work. Opposite intent: waste-on-purpose versus a useful answer.
Cost as signal is the part I keep. If a step is free, I cannot tell whether anyone ran it. If a step has a bill, I can ask what the bill purchased. That question is legitimate in both pictures. It does not make an LLM stack a consensus protocol.
The failure modes split the same way. A bad proof of work is an invalid ticket, so the log does not move. A bad inference can still be a valid decode: the model produced those tokens, and the gate then says they do not do the job. I paid for a candidate, and the candidate failed.
Useful-PoW is a concept, not a protocol
There is a bridge idea, and I want it named so it does not sneak in as a product.
Useful-PoW is the thought that the same checkable spend could also do a job someone wanted anyway: verifiable work that is not pure waste. It remains a concept. It is not a product I ship, and it is not a name I put on xllama, mklang, or ragfs.
Two limits stay in place:
- A useful side effect does not invent a consensus rule. If I cannot say what the ticket admits, who verifies it, and what happens when two valid tickets conflict, I have a workload, not a protocol.
- Inference alone is not a consensus protocol. A model plus a prompt can be costly and checkable and still say nothing about who appends to a shared log. An eval or a price tag does not close that gap.
I use the bridge to keep two pictures in view, not to weld them. The calculator analogy already showed that a rhyme can be exact in one place and empty in another. Cost is the exact place. Identity is the empty one. I will not flatten this into "we reinvented proof of work."
Practice: cost, ledger, and retrieval honesty
The practice I can stand behind is narrower than a Useful-PoW pitch. I keep three honest books in public artifacts: what the forward pass cost, what I already ordered when a gate refuses the output, and what I actually retrieved.
xllama: measure the cost gate, do not rename it
xllama is my local inference project on Xbox developer-mode hardware: GGUF and ONNX Runtime GenAI, CPU or GPU chosen per workload, no retail store path. The Xbox Series S report is the narrative receipt. The generated benchmark summary is the command-level receipt, each figure tied to the command that produced it. I do not repeat a tok/s number here.
What I take from that work is the gate, not a slogan. Local inference has a bill in watts, memory, and wall time. I can measure it and publish the command. I cannot turn that bill into proof of work by calling the measurement a puzzle.
Xbox is a Microsoft trademark. xllama is an independent research project, not a Microsoft product and not affiliated with Microsoft.
mklang: a reject does not erase the tokens I ordered
mklang is a declarative DSL for LLM-driven state machines. I use it on this site as a selective runtime for propose-and-repair loops; the accepted decision is ADR-0003. The machine can halt, repair, or escalate. Deterministic gates still decide. The live validation log records input and output tokens on those runs, including a copy-review rewrite later marked corrected.
That is ledger honesty, not a charge product. Refusing bad output does not erase the tokens I already ordered. The spend happened before the verdict. The mklang project card is the public index for the repository. I am not calling mklang a Useful-PoW product.
ragfs: retrieval has a bill; receipts beat "trust me"
RAGFS mounts an indexed directory and exposes operations, results, and undo through the filesystem. I wrote the interface decision in Why I Gave Agents a Filesystem Instead of Another API. Search on that mount runs a local embedding model and stores vectors in LanceDB. The first query pays a download and an index; later queries pay the path the public documentation describes. I have not published a retrieval-quality benchmark for this setup, so I do not invent one here.
Retrieval is not a free preface to inference. It has a bill, and it has a result I can read: which path I asked for, what the JSON said, whether an undo_id exists. That is a receipt. "Trust me, the context was relevant" is not. The diagnostic order I use when retrieval is weak remains RAG in Production. None of that is proof of work. It is scarce, verifiable work with the opposite intent.
Cost rhymes; equivalence doesn't
Cost rhymes. Equivalence does not.
Proof of work spends compute against a rule so that consensus can stay scarce. Inference spends compute against a model and a prompt so that someone can use the output. Useful-PoW is the concept that those two pictures might one day share a workload. It is not a name for a local runtime, a copy-review machine, or a FUSE mount.
The related posts I want next to this one are the receipts, not a scene: the Xbox cost gate, the filesystem interface for retrieval, the RAG diagnostic order, and the eval harness. Read them as cost, ledger, and retrieval honesty.
FAQ
Is inference a form of proof of work?
No. Inference spends compute against a model and a prompt for a useful answer. Proof of work spends compute against a rule for scarce consensus, and the work is deliberately useless. They share a rhyme. They are not the same protocol.
What does "cost as signal" mean here?
If a step is free, I cannot tell whether anyone ran it. If a step has a bill, I can ask what it purchased and whether a later gate accepted the result. That signal does not make inference into proof of work.
What is Useful-PoW in this article?
A concept only: verifiable work that also does a job someone wanted. It is not a product I ship, and not a name I put on xllama, mklang, or ragfs. Inference alone is still not a consensus protocol.
Why mention xllama, mklang, and ragfs?
They are public receipts for three honesty habits: measure the local cost gate, record tokens already ordered when a gate refuses output, and keep retrieval results instead of "trust me" context.
Does a failed eval erase the compute I spent?
No. The forward pass happens before the gate. If the eval fails, the cost was real and the answer was not. That is ledger honesty, not a proof-of-work ticket.
Related articles
A Refused Answer Still Has to Charge Its Tokens
I keep the tokens that a produce call already billed when the runtime refuses the answer. Ledger honesty in governable agents, from mklang 1.3.7 / pull 111.
Oct 4, 20268 min read#mklang#LLM#Cost#Agents#ProductionRAG in Production: Fix Chunking and Re-Ranking Before Touching Embeddings
When retrieval is weak, swapping embeddings rarely fixes it. Diagnose chunking and re-ranking first.
Dec 20, 202412 min read#RAG#Retrieval#LLM#ProductionFive Function-Calling Patterns I Use in Production
Five function-calling patterns I use in production systems, and the anti-patterns they replaced.
Nov 25, 202413 min read#LLM#Function Calling#OpenAI#Production