Home
8 min read
#mklang#LLM#Cost#Agents#Production

A Refused Answer Still Has to Charge Its Tokens

I keep the tokens that a produce call already billed when the runtime refuses the answer. Ledger honesty in governable agents, from mklang 1.3.7 / pull 111.

Why this matters

I cannot govern an agent if a refused answer disappears from the ledger. The provider already billed the call. If RunResult.usage and the step cost read 0, I lie to myself about what the run spent.

This is bookkeeping, not a faster halt. A truncated answer under on_truncate: halt, or a parse: failure, is still a produce that happened. The policy can refuse the text. It cannot pretend the tokens were never ordered.

The language cut lives in mklang: the document is the program. This note is the ledger cut. The public receipt is gianlucamazza/mklang#111, shipped as package 1.3.7 on 2026-09-29. I am reading that tree. I am not adding a benchmark.

What the receipt records

CHANGELOG 1.3.7 states the bug in one sentence: a produced answer the run refuses was still charged after the fix, and was not charged before it.

Before the fix, deps.llm.produce(...) could return a truncated answer or text that failed parse:. The engine raised a bare ValueError. The run loop's generic except Exception turned that into state-error: … and dropped the tokens of the call that had already been made and billed. The changelog records the production sighting: a ~4096-token answer under on_truncate="halt". That number is the receipt's own wording. It is not a throughput claim.

The halt reason was already correct. The ledger was not.

OutputRejected is the typed error that closes the gap. It carries the call's tokens the way CallFailed already did for a sub-machine:

class OutputRejected(MklangError):
    """A produce call answered (and was billed) but its answer was refused."""

    def __init__(self, error: str, input_tokens: int = 0, output_tokens: int = 0):
        super().__init__(error)
        self.error = error
        self.input_tokens = input_tokens
        self.output_tokens = output_tokens

_exec_produce raises it for the truncation halt and for parse failures. The single-state path catches it before the generic handler, charges the step, records it with its cost, and halts with the unchanged reason: state-error: output-truncated, state-error: parse-list-truncated, or state-error: parse-json: …. A fan-out branch refused the same way keeps its tokens in the [branch-error: …] marker, as a failed call branch already did.

The project card is the public index. Spec 0.4 and package 1.3.7 are the versions the tree declares. I do not treat that as a maturity certificate.

What used to go missing

The missing charge was a control-flow accident. After produce returned, two refusal paths shared one bad raise:

  1. Truncation halt. on_truncate: halt saw Produced.truncated and raised ValueError("output-truncated").
  2. Parse failure. Partial JSON from a length stop, or text that was not valid JSON at all, raised ValueError from _parse_structured.

Both are refusals of an answer that already existed. The generic handler did not know they had been billed. CallFailed already knew: a halted sub-machine keeps input_tokens and output_tokens so the parent can charge them. Refused output did not have that type, so the tokens died in except Exception.

That is the kind of hole I care about in a governable agent. The machine can halt. The machine can repair. The machine cannot lose the bill for work it already ordered and then tell me the run was cheap.

I am not claiming the old path was rare. I am claiming it was dishonest when it happened. The changelog names one production case. The tests name the contract.

What the halt must keep

Two facts have to survive together: the reason, and the cost.

tests/engine/test_truncation.py is the fixture. Four tests failed on usage == 0 before the fix:

  • test_halt_on_truncation_still_charges_the_call — halt reason state-error: output-truncated; usage and the halted step's cost keep the call's tokens.
  • test_parse_truncated_halt_still_charges_the_call — state-error: parse-list-truncated; same charge.
  • test_unparseable_output_halt_still_charges_the_call — state-error: parse-json: output is not valid JSON (; same charge.
  • test_refused_branch_keeps_its_tokens — fan-out sample: 2 finishes done with [branch-error: parse-json: …] markers; both branches' tokens stay on the step.

The fixture LLM returns input_tokens=1200 and output_tokens=4096. Those are test values, not a measured production distribution. The two-branch case doubles them because two produces ran. I am not publishing a cost model from that.

r = run(m, {}, {m.name: m}, _billed("cut", truncated=True), TIERS, on_truncate="halt")
assert (r.status, r.error) == ("halt", "state-error: output-truncated")
assert r.usage == {"input_tokens": 1200, "output_tokens": 4096}
assert r.trace[-1]["cost"] == {"input_tokens": 1200, "output_tokens": 4096}

The engine comment in src/mklang/engine.py is the policy I want on the wall: from the moment produce returns, the call was billed; a refused answer halts via OutputRejected, which carries the tokens so the halt still charges them.

_halt_output_rejected charges, writes policy="state-error", records the step, and returns _halt(f"state-error: {e.error}"). The string a dashboard already keyed on does not change. The ledger does.

What I did not change

The pull request is explicit about the boundary.

Judge calls stay out. Adapters raise JudgeUnparseable before exposing usage. That is a different path. I did not fold it into OutputRejected. A green 1.3.7 install does not mean a failed gate-judge now charges the same way. If I need that later, it is another receipt.

Halt reasons stay the same. state-error: output-truncated is still state-error: output-truncated. I did not invent a new status to celebrate honesty. The point is that the old reason now sits next to a non-zero cost.

No performance, safety, or maturity claim. Charging a refused call does not make the model better, the gate wiser, or the package production-ready. It makes the book true. Bank-grade evals still apply after: a charged halt is not a proof the policy was right.

I also do not treat this as a consensus trick or a market signal. Tokens on a refused step are a bill. They are not a ticket.

Practice: keep the bill next to the refuse

The practice I can stand behind is narrower than a cost-product pitch.

When I run a machine that may halt on truncate or parse, I want three things visible on that step: the halt reason, the input tokens, and the output tokens. If usage is 0 after a produce that returned text, the ledger is lying. mklang 1.3.7 is the interpreter change that stops that lie for those two refusal paths. I do not drop the package into a host and call the host governed.

If a fan-out branch is refused, I want its tokens in the step the same way a failed call already kept them. If a judge fails to parse a choice, I do not pretend 1.3.7 covered it. If I cannot say those things, I cannot account for the run.

The mklang article remains the place I describe the document. This page is the place I describe the bill that survives a refuse.

FAQ

Why charge an answer the runtime refuses?

Because the produce already happened and the provider already billed it. Refusing the text is a policy decision. Erasing the tokens is a lie about cost. I cannot govern a run whose ledger drops work it ordered.

Which halt reasons stay the same?

state-error: output-truncated, state-error: parse-list-truncated, and state-error: parse-json: …. The pull request keeps those strings. Only the charge changes.

Do judge calls charge the same way?

No. The receipt leaves judge calls as they were. Adapters raise JudgeUnparseable before exposing usage. That is not the OutputRejected pattern.

Is this a performance or security win?

No. It is ledger honesty for two refusal paths in the reference interpreter. I am not claiming a faster halt, a safer agent, or a maturity upgrade.

Where is the receipt?

mklang#111, CHANGELOG 1.3.7, errors.py, and tests/engine/test_truncation.py. The project card is the public index.

Share this article

Related articles

  • mklang: The Document Is the Program

    A declarative .mkl file for LLM-driven state machines. Models generate. The machine decides what happens next. What the language guarantees and does not.

    Oct 2, 20267 min read
    #mklang#LLM#DSL#State-Machine#Agents
  • Proof of Work vs Inference: Cost as Signal, Not Identity

    Inference is not proof of work. I separate two costly, checkable spends of compute: a rule for scarce consensus, and a model plus prompt for utility.

    Oct 2, 202610 min read
    #Inference#Cost#Eval#LLM#Production
  • Untrusted Data Must Not Own Control Flow

    I show how reasoning-kernel keeps untrusted data off the effect path: two untrusted reasoners, one deterministic gate. A topology, not a safety certificate.

    Oct 2, 20268 min read
    #Agents#Security#LLM#Python#Production