# Why I Gave Agents a Filesystem Instead of Another API

> A concrete RAGFS workflow shows why I put file operations, results, and undo on a Linux FUSE mount.

Published: 2026-09-25
Canonical: https://gianlucamazza.it/en/blog/agent-filesystem
Tags: FUSE, RAG, Rust, Linux

## Why this matters

An agent working in a directory can already list, read, and write paths. Giving it a separate API for file management creates a second interface for the same material. I wanted an operation's result, including [the way to undo it](/en/blog/ai-workflow-engines), to be available through the filesystem too.

[RAGFS](https://github.com/Venere-Labs/ragfs) is my Linux FUSE implementation. It mounts an indexed directory and exposes a virtual `.ragfs` control directory. The underlying files remain in the source directory. The [ragfs card](/en/projects#ragfs) lists the repository and the documentation.

## One operation, from request to undo

Suppose an agent decides that `docs/old.md` should be removed. The README quick start keeps the mount in the foreground:

```plaintext
mkdir ~/ragfs-mount
ragfs mount ~/Documents ~/ragfs-mount --foreground
```

The delete example then runs against that mount. It writes a relative path to the delete control file and reads the JSON result:

```plaintext
echo "docs/old.md" > ~/ragfs-mount/.ragfs/.ops/.delete
cat ~/ragfs-mount/.ragfs/.ops/.result
```

The delete operation moves the file to trash. [The result contains an `undo_id`](/en/blog/memory-architectures). An agent can retain that identifier and write it to the safety control file to undo the operation:

```plaintext
echo "<undo_id>" > ~/ragfs-mount/.ragfs/.safety/.undo
```

The important part is the sequence: request, inspect the result, keep the undo identifier, and undo the operation if needed. [This is an interface contract](/en/blog/llm-function-calling-patterns), not evidence that an agent will always make the right deletion decision. I still need to review what an agent intends to remove.

## Why search belongs on the same mount

The same mount can search the directory. RAGFS extracts text, splits it into chunks, runs `thenlper/gte-small` locally through Candle, and stores vectors in LanceDB. The model downloads on first use. Later embedding queries use that local model and do not call an external embedding API. The LanceDB index stays outside the source tree. The search path sits on the mount, next to the files.

This is a practical index, with limits. Code chunking detects function and class signatures by pattern matching. The README describes that split as regex on signatures, not a tree-sitter AST. Cosine search stays an exact scan until an IVF-PQ index is built. The README says it builds an IVF-PQ index at 256 chunks or more. L2 and dot-product search stay exact. I have not published a retrieval-quality or scale benchmark for this setup, so I do not treat those implementation choices as proof of better answers.

## What I would not delegate yet

The README marks the CLI, the FUSE mount, agent operations under `.ops/`, and the safety layer under `.safety/` as stable. Organize, dedupe, and cleanup are beta: they propose a plan that needs approval before execution. Python bindings and the MCP server are beta too. [A file interface does not remove the need for that approval boundary](/en/blog/ai-workflow-engines).

FUSE currently runs on Linux. The [repository README](https://github.com/Venere-Labs/ragfs#feature-status) records the status of each surface, installation requirements, and supported extractors. I would check it before adopting RAGFS for a new corpus. This article describes the interface decision, not a production readiness guarantee.

## Where retrieval fits

The diagnostic order I use when a RAG pipeline performs poorly lives in [RAG in Production](/en/blog/rag-systems-production): measure, then inspect chunking and ranking before changing the embedding model. RAGFS gives an agent a way to work with the corpus. It does not establish that the retrieval results are good [without an evaluation](/en/blog/bank-grade-agent-evals).

## FAQ

### Why use a filesystem interface for agent file operations?

I use the filesystem because an agent can already list, read, and write paths there. A separate file-management API would create a second interface for the same material, whereas the filesystem can expose both an operation result and its undo.

### How does RAGFS delete and undo a file operation?

I write a relative path to the delete control file and read the JSON result. The delete operation moves the file to trash and returns an undo_id. I can retain that identifier and write it to the safety control file to undo the operation.

### Does RAGFS search call an external embedding API?

No. I run thenlper/gte-small locally through Candle, and the model downloads on first use. Later embedding queries use that local model rather than an external embedding API; LanceDB stores vectors outside the source tree.

### What are the current retrieval limits in RAGFS?

I use pattern matching for code function and class signatures rather than a tree-sitter AST. Cosine search is an exact scan until an IVF-PQ index is built, with the README saying it builds an IVF-PQ index at 256 chunks or more. I have not published retrieval-quality or scale benchmarks.

### Which RAGFS features should still require approval?

I treat organize, dedupe, and cleanup as beta because they propose a plan that needs approval before execution. Python bindings and the MCP server are also beta. A file interface does not remove the need for an approval boundary.
