mklang: The Document Is the Program
A declarative .mkl file for LLM-driven state machines. Models generate. The machine decides what happens next. What the language guarantees and does not.
A .mkl file describes an agent as a state machine. An LLM executes the generative steps. The document is the program; the host supplies the interpreter, the tools, and the code-hook gates.
Models generate. Machines decide what happens next.
The project card is on Projects. The repository is gianlucamazza/mklang, Apache-2.0. Language spec 0.4. Reference package 1.3.7. Docs at docs.mklang.dev.
Why this matters
I have shipped agent control flow in Python. After a few months the graph is real, but it lives in constructors, closures, and whichever provider the process was configured with. A diff on the behavior is a diff on application code. That is fine for an application. It is a bad artifact when the thing I want to inspect, version, and hand over is the control flow itself.
LangGraph is the right tool when the graph is application code. I still use that shape. mklang is the other side of the same problem: a portable document, interpreted by an LLM runtime, with the topology written down.
mklang : LangGraph :: a declarative spec : Python code
The trade is explicit. You get a provider-independent program and a control flow you can read without the interpreter. You give up host-language expressiveness. Production machines still need a developer for tools, hooks, budgets, and untrusted inputs.
I am not claiming adoption, a benchmark, or that two models will take the same gate. The useful result is the constraint: a diff on the .mkl is a diff on the machine.
What a machine is
Four entities.
- machine — one
.mklfile. Entry state, global budget, map of states. - state — where something happens.
- context — a blackboard that accumulates across the run. The only memory channel between states. A state's output is deposited under a key; prompts read it with
{{…}}. - tier —
fast,balanced, orreasoning. Capability, not vendor.
Pinning a provider or a model inside the document is a deliberate non-goal. It would break portability. The host maps tiers to concrete models in runtime.yaml. The same file runs on DeepSeek, Anthropic, OpenAI, Google, OpenRouter, xAI, Mistral, or a keyless local endpoint such as Ollama.
Each generative state has four faces.
| Face | Question | Rule |
| ----------- | -------------- | ------------------------------------------------------ |
| structure | What shape? | Prose, not a type |
| prompt | What to think? | Task, with {{…}} interpolation |
| execution | How to act? | Sticky policy. Never a side effect |
| gates | When to exit? | Natural-language conditions. These are the transitions |
A gate resolves to ok, repair, escalate, or fail, then routes. Side effects are not prose. A tool: state calls a host callable; the observation re-enters the context. Optional faces cover the patterns I actually reach for: reason for a traced chain of thought, accumulate for list-append, sample / over for fan-out, call for another machine. A step budget stops a runaway loop. Checkpoints are resumable.
The smallest machine:
machine: greet
entry: answer
states:
answer:
prompt: "Greet the user in one sentence."
output: reply
gates:
- when: the reply is a greeting
then: ok
to: END
The model produces reply. A judge decides whether it is a greeting. The topology is in the file, not in a system prompt.
The spec's pattern cookbook maps the usual agent shapes onto this core: chain-of-thought, ReAct, Reflexion, self-consistency, tree-of-thought, plan-and-execute, debate, map-reduce, router, speculative cascade. examples/ has react.mkl, triage.mkl, and self_consistency.mkl. A small standard library (std_self_consistency, std_refine, and others) is callable from the CLI or the console.
What the host does
The document does not compile. A conformant runtime is any host with access to an LLM. The reference interpreter is Python.
pipx install 'mklang[mcp]'
mklang init --user
mklang console
init --user scaffolds config and .env under ~/.config/mklang/, and machines under ~/.local/share/mklang/machines/, without overwriting. The console is the front door: an agent-first TUI whose brain is itself a machine, agent.mkl, with no privileged powers. Sessions persist (--continue). Escalations and tool consent are interactive. The default brain reads the workspace through bounded read-only tools. No shell, no write.
mklang check validates. mklang doctor reports which config layer won, which keys are missing, and which tool backends are live. mklang test runs a machine against named scenarios with a scripted LLM, so a control-flow fixture does not need an API key. The conformance suite is implementation-neutral YAML. A second runtime, in TypeScript or Rust, conforms if it passes the same cases. Conformance is the mechanical contract. It is not a promise that two models will judge a gate the same way.
Editor validation hangs off the published JSON Schema:
https://raw.githubusercontent.com/gianlucamazza/mklang/main/schema/mklang.schema.json
What it is not
The spec lists this on purpose.
- It does not compile to a formal artifact.
- It does not guarantee determinism.
- It has no static types for
structureor for gate conditions. Both are prose, judged at runtime. - Gate accuracy and cross-provider stability are empirical. I have seen gates diverge across providers. That is a measurement, not a bug in the document.
Untrusted context is delimited, not judged. Host inputs, tool observations, and deposits are tainted. Tainted interpolations are fenced as <data-NONCE>…</data-NONCE> with a fresh nonce per call, and the model is told that fenced content is not an instruction. A judge decision over external data marks the transition. An effectful tool: state reached on that path can be refused with --untrusted-flow halt. mklang lint names those states at authoring time.
Exact policy does not belong on a prose gate. Amounts and allowlists go on a hook: gate, a host predicate with no model in the path, or on an escalation before an irreversible effect. Sandboxed tool brokers, signed context zones, and non-interference proofs are explicit non-goals.
This repository is the language and the reference interpreter. A hosted platform, if it exists later, is a separate thing. Do not treat the Apache-2.0 tree as an orchestrator you can drop into production and forget.
Where it sits next to the other work
LangGraph remains the place where state is an API: channels, reducers, checkpoints of state rather than of side effects. mklang does not replace that. It is the document I want when the machine should outlive the Python that happens to run it.
orka is the other boundary: a Rust runtime with hard capability limits around multi-agent work. mklang is the inspectable program. orka is the host that should make an unauthorized effect hard. They are not the same layer, and I do not pretend the .mkl file is a capability system.
FAQ
Is mklang a replacement for LangGraph?
No. LangGraph is application code that builds and runs a graph. mklang is a portable document interpreted by an LLM runtime. Use LangGraph when you need host-language expressiveness. Use a .mkl when the control flow itself is the artifact you want to diff.
Does a green conformance run mean the agent is correct?
No. Conformance pins mechanical semantics: interpolation, routing, budgets, taint marking. Gate outcomes depend on the model. Two conformant runtimes can diverge in production.
Where do side effects go?
In a tool: state, which is a host callable, or behind a hook: gate for exact policy. Never in execution, and never as a prose instruction the model is asked to confirm.
What should I read first?
The project card, then SPEC.md, then hello.mkl from mklang init --user. The console guide is docs/guides/console.md in the repo.
Related articles
A Refused Answer Still Has to Charge Its Tokens
I keep the tokens that a produce call already billed when the runtime refuses the answer. Ledger honesty in governable agents, from mklang 1.3.7 / pull 111.
Oct 4, 20268 min read#mklang#LLM#Cost#Agents#ProductionUntrusted Data Must Not Own Control Flow
I show how reasoning-kernel keeps untrusted data off the effect path: two untrusted reasoners, one deterministic gate. A topology, not a safety certificate.
Oct 2, 20268 min read#Agents#Security#LLM#Python#ProductionIs Saying an LLM Doesn't Think Like Saying a Calculator Can't Do Numbers?
Where the calculator analogy for LLMs holds and where it breaks: what interpretability, chain-of-thought and philosophy of mind say about thinking.
Jul 2, 202616 min read#LLM#AI Reasoning#Interpretability#Philosophy of Mind