# mklang: The Document Is the Program

> A declarative .mkl file for LLM-driven state machines. Models generate. The machine decides what happens next. What the language guarantees and does not.

Published: 2026-10-02
Canonical: https://gianlucamazza.it/en/blog/mklang
Tags: mklang, LLM, DSL, State-Machine, Agents

A `.mkl` file describes an agent as a state machine. An LLM executes the generative steps. The document is the program; the host supplies the interpreter, the tools, and the code-hook gates.

Models generate. Machines decide what happens next.

The project card is on [Projects](/en/projects#mklang). The repository is [gianlucamazza/mklang](https://github.com/gianlucamazza/mklang), Apache-2.0. Language spec 0.4. Reference package 1.3.7. Docs at [docs.mklang.dev](https://docs.mklang.dev/).

## Why this matters

I have shipped agent control flow in Python. After a few months the graph is real, but it lives in constructors, closures, and whichever provider the process was configured with. A diff on the behavior is a diff on application code. That is fine for an application. It is a bad artifact when the thing I want to inspect, version, and hand over is the control flow itself.

[LangGraph](/en/blog/langgraph-workflow-orchestration) is the right tool when the graph is application code. I still use that shape. mklang is the other side of the same problem: a portable document, interpreted by an LLM runtime, with the topology written down.

```plaintext
mklang : LangGraph :: a declarative spec : Python code
```

The trade is explicit. You get a provider-independent program and a control flow you can read without the interpreter. You give up host-language expressiveness. Production machines still need a developer for tools, hooks, budgets, and untrusted inputs.

I am not claiming adoption, a benchmark, or that two models will take the same gate. The useful result is the constraint: a diff on the `.mkl` is a diff on the machine.

## What a machine is

Four entities.

- **machine** — one `.mkl` file. Entry state, global budget, map of states.
- **state** — where something happens.
- **context** — [a blackboard that accumulates across](/en/blog/memory-architectures) the run. The only memory channel between states. A state's output is deposited under a key; prompts read it with `{{…}}`.
- **tier** — `fast`, `balanced`, or `reasoning`. Capability, not vendor.

Pinning a provider or a model inside the document is a deliberate non-goal. It would break portability. The host maps tiers to concrete models in `runtime.yaml`. The same file runs on DeepSeek, Anthropic, OpenAI, Google, OpenRouter, xAI, Mistral, or a keyless local endpoint such as Ollama.

Each generative state has four faces.

| Face        | Question       | Rule                                                   |
| ----------- | -------------- | ------------------------------------------------------ |
| `structure` | What shape?    | Prose, not a type                                      |
| `prompt`    | What to think? | Task, with `{{…}}` interpolation                       |
| `execution` | How to act?    | Sticky policy. Never a side effect                     |
| `gates`     | When to exit?  | Natural-language conditions. These are the transitions |

A gate resolves to `ok`, `repair`, `escalate`, or `fail`, then routes. Side effects are not prose. A `tool:` state calls a host callable; the observation re-enters the context. Optional faces cover the patterns I actually reach for: `reason` for a traced chain of thought, `accumulate` for list-append, `sample` / `over` for fan-out, `call` for another machine. A step budget stops a runaway loop. [Checkpoints are resumable](/en/blog/ai-workflow-engines).

The smallest machine:

```plaintext
machine: greet
entry: answer
states:
  answer:
    prompt: "Greet the user in one sentence."
    output: reply
    gates:
      - when: the reply is a greeting
        then: ok
        to: END
```

The model produces `reply`. A judge decides whether it is a greeting. The topology is in the file, not in a system prompt.

The spec's pattern cookbook maps the usual agent shapes onto this core: chain-of-thought, ReAct, Reflexion, self-consistency, tree-of-thought, plan-and-execute, debate, map-reduce, router, speculative cascade. `examples/` has `react.mkl`, `triage.mkl`, and `self_consistency.mkl`. A small standard library (`std_self_consistency`, `std_refine`, and others) is callable from the CLI or the console.

## What the host does

The document does not compile. A conformant runtime is any host with access to an LLM. The reference interpreter is Python.

```plaintext
pipx install 'mklang[mcp]'
mklang init --user
mklang console
```

`init --user` scaffolds config and `.env` under `~/.config/mklang/`, and machines under `~/.local/share/mklang/machines/`, without overwriting. The console is the front door: an agent-first TUI whose brain is itself a machine, `agent.mkl`, with no privileged powers. Sessions persist (`--continue`). Escalations and tool consent are interactive. The default brain reads the workspace through bounded read-only tools. No shell, no write.

`mklang check` validates. `mklang doctor` reports which config layer won, which keys are missing, and which tool backends are live. `mklang test` [runs a machine against named scenarios](/en/blog/bank-grade-agent-evals) with a scripted LLM, so a control-flow fixture does not need an API key. The conformance suite is implementation-neutral YAML. A second runtime, in TypeScript or Rust, conforms if it passes the same cases. Conformance is the mechanical contract. It is not a promise that two models will judge a gate the same way.

Editor validation hangs off the published JSON Schema:

`https://raw.githubusercontent.com/gianlucamazza/mklang/main/schema/mklang.schema.json`

## What it is not

The spec lists this on purpose.

- It does not compile to a formal artifact.
- It does not guarantee determinism.
- It has no static types for `structure` or for gate conditions. Both are prose, judged at runtime.
- Gate accuracy and cross-provider stability are empirical. I have seen gates diverge across providers. That is a measurement, not a bug in the document.

Untrusted context is delimited, not judged. Host inputs, tool observations, and deposits are tainted. Tainted interpolations are fenced as `<data-NONCE>…</data-NONCE>` with a fresh nonce per call, and the model is told that fenced content is not an instruction. A judge decision over external data marks the transition. An effectful `tool:` state reached on that path can be refused with `--untrusted-flow halt`. `mklang lint` names those states at authoring time.

Exact policy does not belong on a prose gate. Amounts and allowlists go on a `hook:` gate, a host predicate with no model in the path, or on an escalation before an irreversible effect. Sandboxed tool brokers, signed context zones, and non-interference proofs are explicit non-goals.

This repository is the language and the reference interpreter. A hosted platform, if it exists later, is a separate thing. Do not treat the Apache-2.0 tree as an orchestrator you can drop into production and forget.

## Where it sits next to the other work

[LangGraph](/en/blog/langgraph-workflow-orchestration) remains the place where state is an API: channels, reducers, checkpoints of state rather than of side effects. mklang does not replace that. It is the document I want when the machine should outlive the Python that happens to run it.

[orka](https://github.com/gianlucamazza/orka) is the other boundary: a Rust runtime with hard capability limits around multi-agent work. mklang is the inspectable program. orka is the host that should make an unauthorized effect hard. They are not the same layer, and I do not pretend the `.mkl` file is a capability system.

## FAQ

### Is mklang a replacement for LangGraph?

No. LangGraph is application code that builds and runs a graph. mklang is a portable document interpreted by an LLM runtime. Use LangGraph when you need host-language expressiveness. Use a `.mkl` when the control flow itself is the artifact you want to diff.

### Does a green conformance run mean the agent is correct?

No. Conformance pins mechanical semantics: interpolation, routing, budgets, taint marking. Gate outcomes depend on the model. Two conformant runtimes can diverge in production.

### Where do side effects go?

In a `tool:` state, which is a host callable, or behind a `hook:` gate for exact policy. Never in `execution`, and never as a prose instruction the model is asked to confirm.

### What should I read first?

The [project card](/en/projects#mklang), then [SPEC.md](https://github.com/gianlucamazza/mklang/blob/main/SPEC.md), then `hello.mkl` from `mklang init --user`. The console guide is `docs/guides/console.md` in the repo.
