Research and demos
emotional-memory
Memory layer for LLM systems with affective state encoding, a PyPI package, Zenodo DOI, and reproducible benchmarks against Mem0, LangMem, and Letta.
Open the artifactInspect emotional-memoryI select a few projects here to show different engineering problems. The complete archive follows, with released software, research and demos, and system notes kept distinct.
Start with the selected work, then inspect the evidence and technical details in the archive. Research prominence does not imply production deployment.
Research and demos
Memory layer for LLM systems with affective state encoding, a PyPI package, Zenodo DOI, and reproducible benchmarks against Mem0, LangMem, and Letta.
Open the artifactInspect emotional-memoryShipped software
Rust agent runtime that takes work from chat and HTTP channels, assigns priority, and routes it into bounded LLM workflows with MCP/A2A support.
Inspect orkaResearch and demos
A Python reference implementation of a reasoning kernel that separates untrusted model text from authorized effects using capability-based control and taint tracking.
Inspect reasoning-kernelShipped software
Shipped local-first Python system: durable workflow state, event audit trail, recovery before external effects, and inspectable assistant tools.
Inspect openfattureAll 20 projects remain available below. Research repositories and non-LLM pieces are labelled separately. They are not production work.
active
Shipped local-first Python system: durable workflow state, event audit trail, recovery before external effects, and inspectable assistant tools.
How do you persist workflow state, recover a failed export, and keep assistant actions inspectable when those actions can trigger irreversible external effects?
A shipped local-first Python system: persisted workflow state, an event audit trail, recovery before export or send, and inspectable assistant tools.
A public package on GitHub and PyPI. The dated Phase 2 test report is a historical snapshot, not a current quality score.
The Phase 2 test report is dated historical evidence, not a current quality score. External effects stay behind manual review; inspectability outranks autonomous action.
Local-first operation limits convenience but improves control over workflow state and side effects.
Expand workflow checks, recovery coverage, and auditable AI assistance.
active
Rust agent runtime that takes work from chat and HTTP channels, assigns priority, and routes it into bounded LLM workflows with MCP/A2A support.
Requests arrive from chat, email, and internal tools, but without a single durable queue there is no reliable path from channel input to a tracked, reviewable LLM workflow.
An agent runtime that turns channel input into prioritized LLM workflows with MCP/A2A protocol adapters, sandboxed execution paths, and explicit runtime boundaries.
One prioritized queue routes multi-channel requests into bounded LLM workflows with MCP/A2A support.
The runtime must stay protocol-oriented and avoid coupling to a single interface.
A lower-level runtime gives more control but requires sharper product boundaries.
Harden observability, workspace policy, and durable execution semantics.
active
Secure transport layer for agent messages, with end-to-end encryption, DID identity, relay delivery, MCP adapters, and A2A interoperability.
Agent systems need a transport layer for cross-boundary messages without shared keys, central trust, or relay access to plaintext.
A protocol layer for agent-to-agent messages using W3C DID identity, end-to-end encryption, relay delivery, MCP connector support, and A2A interoperability.
A secure transport pattern for agent integrations: the relay can deliver and queue messages, but message content remains outside its trust boundary.
The relay routes messages but must not become the trust anchor for identity or message confidentiality.
Cryptographic ownership increases trust clarity while adding connector and key-management complexity.
Broaden connector distribution, production billing, and interoperability paths.
active
Native Chromecast sender work for LibreWolf and Firefox, with an openscreen backend for media casting and Wayland screen mirroring without third-party desktop tools.
Browser casting on Linux often depends on external desktop tools or incomplete paths, especially when screen mirroring enters the workflow.
A native Cast sender path for LibreWolf and Firefox backed by openscreen work, with a focus on media casting and Wayland screen mirroring.
A systems-level showcase for browser integration work: protocol boundaries, native media paths, and desktop constraints are handled below the web UI layer.
The work is platform-sensitive: display server, codec, browser, and Cast receiver behavior all shape the implementation.
A native path gives better control than a wrapper, but it exposes lower-level compatibility and maintenance work.
Harden receiver compatibility, capture reliability, and browser packaging paths.
active
Arch Linux package for official llama.cpp Vulkan binaries, with automated upstream release checks, reproducible checksums, build verification, and a smoke test for runtime backend loading.
active
Rust semantic browser for AI agents that need structured access to web pages, page state, and interaction surfaces instead of brittle screenshots alone.
Agents that operate on the web need structured page state and interaction targets instead of relying only on brittle screenshots or raw DOM dumps.
A Rust browser layer that exposes semantic page structure and interaction surfaces so agent runtimes can reason over the page with clearer boundaries.
A browser automation proof point focused on inspectable state: the agent sees a structured interface, not an opaque visual stream.
The browser layer must preserve enough page semantics for agents without pretending that arbitrary web pages are deterministic.
Semantic state is easier to inspect than screenshots, but it requires careful handling of dynamic pages and accessibility gaps.
Connect the page model to eval traces and safer browser-control policies.
active
MCP server for Python development tools used by AI assistants.
I wanted a Python-based MCP server that gives an AI assistant a defined toolbox for Python development. The repository metadata supports a narrow scope: an MCP integration surface for developer tools, not a claim about specific workflows, adoption, or performance.
I keep the system boundary at the Model Context Protocol. Assistant clients connect to the server, request development-oriented tools, and receive responses through the protocol interface. Python is the implementation language. The design focus is where the tool boundary sits: what an assistant can invoke is declared by the server, not improvised inside a chat instruction.
The repository presents a Python MCP server for AI-assisted Python development. From the available metadata, I can state the outcome at repository level only: a concrete protocol-based place to expose Python development tools to compatible assistant clients.
The public metadata does not describe the individual tools, execution model, persistence layer, benchmarks, or production usage. I therefore keep the case study at the integration and system-boundary level.
Using MCP gives a clear interface for assistant-tool interaction, but it also means the useful behavior depends on the tool definitions implemented behind that interface. Keeping the scope narrow reduces ambiguity, at the price of pushing persistence and evaluation onto whoever wires the toolbox into a larger system.
The next technical step I would evaluate is documenting each exposed tool with its input contract, failure modes, and expected side effects. That would make the server easier to test and would create a basis for reproducible evals without making unsupported claims about current behavior.
active
A declarative DSL for LLM-driven state machines.
I wanted a small language boundary for LLM-driven state machines: the state machine should be described as a document, not hidden inside imperative glue code. The repository frames the `.mkl` document as the program and the LLM as the runtime, so the main engineering question is how to make agent-like control flow explicit enough to inspect, edit, and version.
I model the project as a declarative DSL around state-machine structure. The source artifact is a `.mkl` document; Python provides the implementation substrate; the LLM is treated as the runtime component that advances the machine according to the document. This keeps the program representation separate from the execution substrate and makes the repository a place to test language shape, parsing boundaries, and runtime responsibilities.
The repository documents an approach for expressing LLM-driven state machines as source files. I do not claim benchmarks, production adoption, or broad model support from the available metadata. The useful result is the architectural constraint itself: a program can be represented as a `.mkl` document while runtime behavior remains attached to an LLM-backed executor.
I keep the description limited to repository metadata: Python, a declarative DSL, `.mkl` documents, LLM-driven state machines, and the idea that the LLM acts as runtime. I do not infer parser features, execution guarantees, model coverage, or operational metrics that are not stated.
A DSL makes the control structure easier to treat as source, but it also introduces language-design work: syntax, validation, runtime semantics, and error reporting must be defined carefully. Keeping the LLM as runtime preserves flexibility, while making deterministic behavior and recovery semantics something the surrounding system must specify explicitly.
The natural next work is to make the contract between `.mkl` documents and the runtime more explicit: state representation, transition rules, validation behavior, and test fixtures. Until a document can be replayed against a fixture and produce the same transitions twice, the language is a shape, not a guarantee.
active
I provide a TypeScript MCP server for Manus.im integration, with Docker, OAuth, Prometheus, monitoring, and security topics.
I built this repository to connect Model Context Protocol clients with Manus.im through a TypeScript server. The repository metadata identifies deployment, authentication, monitoring, and security as engineering concerns around that integration.
I organize the project as an MCP server in TypeScript with a Manus.im integration. I include Docker as a deployment concern, OAuth as an authentication concern, and Prometheus as a monitoring concern; the available metadata does not document lower-level implementation details.
I publish a repository that defines a TypeScript MCP server focused on Manus.im integration. I do not infer performance, adoption, or operational results from the available metadata.
I limit this case study to the repository metadata. I do not claim specific tools, request flows, persistence behavior, recovery behavior, metrics, or security controls that the metadata does not describe.
I keep the technical description at the integration and concern level rather than reconstructing undocumented internals. This avoids treating topic labels as proof of specific implementation choices.
I would need source-level documentation or operational evidence before documenting concrete MCP tools, OAuth flows, Prometheus metrics, Docker configuration, or security mechanisms.
active
A Python MCP server that exposes DuckDuckGo web search to LLM clients.
I addressed the integration problem of making DuckDuckGo web search available to LLM clients through a Model Context Protocol server.
I implemented the project as a Python MCP server, using the protocol boundary to present DuckDuckGo web search as a tool available to an LLM client.
I provide a repository focused on connecting LLM clients to DuckDuckGo web search through MCP. The available metadata does not document benchmark or deployment results.
I scope the project to DuckDuckGo web search exposed through MCP. The repository metadata does not document caching, authentication, rate handling, or result-ranking behavior.
I use MCP as the integration boundary for LLM clients, which keeps the repository focused on a protocol-based search interface rather than a general-purpose search application.
I would evaluate future changes against documented client requirements and the operational behavior of the DuckDuckGo integration.
active
A Python reference implementation of a reasoning kernel that separates untrusted model text from authorized effects using capability-based control and taint tracking.
LLM agents act on untrusted text, so a prompt injection can trigger tool calls and effects the user never authorized.
A reasoning kernel that separates a privileged planner from quarantined untrusted data and gates every effect behind explicit capabilities and taint tracking.
In this Python reference implementation, a class of prompt-injection effects is constrained by construction, with auditable decisions: research, not a production security guarantee.
Security comes from structure, not model behavior: the planner must never act directly on untrusted content.
Explicit capability mediation adds support code but removes a whole class of injection effects.
Broaden the capability catalog and integrate it with real tool runtimes.
active
Memory layer for LLM systems with affective state encoding, a PyPI package, Zenodo DOI, and reproducible benchmarks against Mem0, LangMem, and Letta.
When a business adds an AI assistant, how do you verify that what it remembers today will still recall correctly after the next model or software update?
A research-focused memory layer with affective state encoding, a PyPI package, Zenodo DOI, benchmark artifacts, and comparisons against existing memory frameworks.
In the public Addendum R result, AFT reached 0.595 LLM-judged answer accuracy versus 0.440 for cosine (Δ +0.155, N=200, p<0.001). The repository also records negative results outside affect-discriminative recall, so I treat this as regime-specific evidence rather than general superiority.
The figures come from the repository’s public Addendum R artifact and apply to affect-discriminative recall with oracle affect labels. Stronger scientific claims require broader external validation.
Research rigor has priority over broad framework compatibility or a larger feature set.
Expand human evaluation, semantic confound tests, and longitudinal memory benchmarks.
active
Interpretable affect source for emotional-memory, built from a reduced Drosophila mushroom-body circuit with persistent mood, approach/avoid decisions, journal replay, and committed benchmarks.
active
Agent-to-agent chat with end-to-end encryption over a self-hosted relay, designed so the relay cannot read message contents.
Agents coordinating across operators need a private channel where the relay cannot read messages or impersonate a participant.
A self-hosted agent-to-agent chat layer with end-to-end encryption, a blind relay, and Double Ratchet session security.
Agent coordination over a private, self-hosted channel: no central party can read or block the messages.
The relay forwards ciphertext only; identity and confidentiality must never depend on trusting the server.
Running your own relay increases operational work but removes a central party that could read or block messages.
Broaden client support, group sessions, and key-recovery flows.
active
Local LLM chat and Stable-Diffusion image generation on Xbox Series S|X in UWP development mode, with ONNX Runtime GenAI and DirectML routed per workload.
Local LLM inference on constrained consumer hardware is limited by memory, runtime APIs, packaging, and platform-specific deployment paths.
An Xbox Series S|X inference app in UWP development mode, with ONNX Runtime GenAI and DirectML routed per workload, to test local model execution inside tight platform limits.
A concrete edge-inference proof point: model runtime, device limits, and packaging constraints are made explicit instead of hidden behind a generic demo.
The project is an experiment, not a product claim: platform restrictions and model size limits define the useful boundary.
A constrained device makes the engineering limits visible, but reduces model choice and deployment flexibility.
Measure new GGUF builds on the console before promoting them to catalogue defaults.
active
A documented LangChain RAG pipeline for comparing OpenAI and HuggingFace embeddings without changing the rest of the retrieval flow.
RAG quality depends on chunking, embeddings, and retrieval choices that are hard to compare without a documented baseline.
A reference RAG pipeline on LangChain that runs OpenAI and HuggingFace embeddings over the same documents and queries for side-by-side comparison.
A documented baseline for comparing retrieval choices before investing in a production pipeline.
It is a teaching baseline, not a production service: clarity and reproducibility come before scale.
A notebook format favors readability over deployment, so production concerns are intentionally out of scope.
Add reranking, evaluation sets, and more embedding backends.
active
Experimental Rust bare-metal kernel for Raspberry Pi 4 with cooperative tasks, EL0 agents, IPC, and W^X MMU
I wanted a small Rust kernel for Raspberry Pi 4 that keeps the core operating-system mechanisms visible: cooperative tasks, EL0 agent execution, IPC, and W^X memory permissions. The repository metadata supports an experimental bare-metal AArch64 scope; I do not claim production use, adoption, or benchmark results.
I structured the project as a no_std Rust bare-metal kernel for AArch64. The design centers on a cooperative task model, an EL0 boundary for agents, IPC primitives, and MMU configuration intended to enforce W^X permissions. I keep the boundaries explicit so that privilege changes, communication paths, and memory permissions can be inspected as kernel mechanisms rather than hidden runtime behavior.
The result is a public repository focused on an experimental Rust kernel for Raspberry Pi 4. Its supported scope is the kernel architecture described in the metadata: bare-metal execution, cooperative scheduling, EL0 agents, IPC, and W^X MMU work. Where the metadata is silent, I treat the status as unknown rather than inferred.
This is a bare-metal Raspberry Pi 4 project, so the design stays close to hardware and does not assume a hosted OS runtime, a standard library, or a normal process model. The repository metadata does not provide test coverage, proof details, deployment status, or benchmark data.
Rust and no_std keep the implementation away from hosted runtime assumptions, but they also make hardware-specific work explicit. Cooperative scheduling is easier to inspect than preemption, while it requires tasks to yield intentionally. W^X and EL0 boundaries support isolation goals at the cost of MMU and context-management complexity.
If I continue the project, I would make the verification story easier to inspect: name the invariants the kernel claims to hold, document how to rerun the checks that establish them, and state what remains unproven.
active
A Python project for a Chromium page agent aimed at coding agents.
I created pagouse around a focused integration problem: coding agents need a way to work with Chromium pages. The repository metadata identifies the project as a Chromium page agent for coding agents, implemented in Python.
I kept the project boundary centered on a page-agent role rather than describing a broader automation platform. Python is the implementation language, while Chromium, CLI, MCP, and accessibility are explicit repository topics.
The public metadata documents a Python repository named pagouse with a defined focus on Chromium page interaction for coding agents. It does not provide benchmark, usage, or deployment data, so I do not infer operational results from it.
I limited this case study to the repository metadata. The available information names the language, project description, and topics, but does not describe internal modules, browser-control methods, test coverage, or release practices.
I describe Chromium, CLI, MCP, and accessibility as repository topics rather than confirmed implementation details. This preserves a useful technical boundary without attributing undocumented behavior to the project.
I would update this case study when the repository documents its page-control model, supported interfaces, test strategy, and evaluation criteria.
active
I expose indexed directories through a Linux FUSE mount, with JSON file operations, undo, and local semantic search. Semantic cleanup still needs plan approval. Python bindings and the MCP server are beta.
active
GitHub profile README covering agent systems and production AI infrastructure work.
I use this profile repository to state the technical areas I work on: verifiable agent systems and production AI infrastructure. The repository metadata also identifies agent orchestration, LLMs, MCP, Rust, and Bitcoin as relevant topics.
I keep the project as a GitHub profile README rather than presenting it as an application or service. Its role is to provide a concise, version-controlled technical entry point for the areas represented by the repository topics.
The repository provides a public, inspectable profile document. Based on the available metadata, I do not infer deployed components, measured results, or operational usage beyond that documentation role.
The available metadata describes a profile README and its topics, but not source layout, integrations, deployment, tests, or evaluation procedures. I therefore limit this case study to the repository’s documented role.
A profile README is easy to inspect and maintain through Git history, but it does not by itself specify executable behavior, interface contracts, or evidence from production operation.
I can extend the README with links to implementation repositories, architecture notes, or reproducible evaluation material when those artifacts are publicly available.
Tell me about the operational problem and the constraints you are working with. A few lines are enough for a first technical assessment.