LLM production from prototype
After the review I implement the agreed production changes in your repository.
Brescia · software
Independent engineer who ships LLM systems from prototype to production (state, recovery, eval, cost).
When a working prototype already exists, I start with a production readiness review. That review does not include code changes.
Published software and reproducible benchmarks: 1 Software DOI 10.5281/zenodo.19972258 20 Projects
I put verification and authorization boundaries around model output, and make state, recovery, evaluation, and cost inspectable.
When a working prototype already exists, the first step is the production readiness review. That review does not include code changes. Production readiness review · €3,500
After the review I implement the agreed production changes in your repository.
I take that production work as a fractional CTO engagement: stack choices, code and vendor review, and a delivery roadmap.
I document how to run and maintain your system, with tests, operating procedures, and the access your team needs.
I describe the problem and what I built, with links to the code and supporting evidence.
#emotional-memory
I encode affective state into a memory store and run a published benchmark; the claim matrix feeds back into encoding. Dataset ingest and the published artifacts stay on the project page — this preview omits them.
When a business adds an AI assistant, how do you verify that what it remembers today will still recall correctly after the next model or software update?
On a published benchmark (Zenodo DOI), the method beat a standard cosine baseline on answer accuracy. Negative results outside that regime are in the repo: I treat this as evidence for that case, not a universal claim.
#orka
Channel input is routed into a scheduler that runs a bounded workflow and writes a checkpoint. An interrupted run resumes from that checkpoint. Protocol adapters stay on the project page — this preview omits them.
Requests arrive from chat, email, and internal tools, but without a single durable queue there is no reliable path from channel input to a tracked, reviewable LLM workflow.
One prioritized queue takes requests from several channels into tracked LLM workflows you can review. (MCP/A2A support: detail on the project page.)
Write to info@gianlucamazza.it with the operational problem and the systems involved. When a working prototype already exists, the first step is the production readiness review. That review does not include code changes.
Systems explains architecture patterns; Research links experiments and their limits. The architecture of this site shows how model proposals pass checks before publication.