Back to work / Case study

Arborist MCP.

A Rust-powered semantic code analysis, symbol indexing, and patch validation engine exposed through a stdio MCP / JSON-RPC gateway for AI agents and editor clients.

RoleRust Core / Python Gateway / MCP Tools
Year2026
PlatformRust / Python / MCP Gateway
Status Late-alpha / early-beta
58
MCP tools
10
Parsed languages
3300+
Rust tests
2900+
Commits

AI coding agents need structured code understanding, not just raw text or repository grep.

Arborist MCP started as an infrastructure layer for the gap between what an LLM can read and what it can safely change: symbol graphs, stable identifiers, patch validation, and indexed context.

The result is a Rust core with PyO3 bindings and a thin Python gateway, so editor clients and AI agents can use structured tools instead of ad hoc code scraping.

A Rust core owns parsing, indexing, validation, while the Python layer only orchestrates the MCP protocol.

  1. R1Rust core

    arborist-core handles Tree-sitter parsing, semantic extraction, symbol modeling, VFS state, patch validation, and SQLite-backed persistent indexes.

  2. R2PyO3 bridge

    arborist-py exposes the core as a native Python module while keeping most performance-sensitive logic inside Rust.

  3. R3MCP gateway

    A stdio JSON-RPC gateway registers 58 structured tools, validates requests, tracks lifecycle, and keeps a legacy arborist/* compatibility layer.

Fail-closed correctness and bounded cooperative budgets are the design spine.

Fail-closed indexingover Best-effort guesswork

Corrupt or foreign databases are rejected instead of degraded, and identifier binding checks reject patches that cannot be resolved to visible symbols.

Cooperative cancellationover Unbounded tool calls

Each tool accepts timeout budgets shared across batched read-only calls, with explicit non-interruptible steps in enumeration.

Gateway disciplineover Fat Python service

A facade as thin as 240 lines hides a pure Rust core, avoiding logic drift between the protocol layer and the engine.

The engine is built to make AI modifications safer by keeping context shallow, verifiable, and replayable.

Symbol graph

Functions, classes, variables, signatures, docstrings, overload identities, and byte ranges are extracted into stable, queryable symbols.

VFS edit stream

Open, change, close, virtual patch, byte edit, commit, discard, and incremental Tree-sitter reparse keep unsaved source behavior consistent.

Patch validation

AST and position patches, previews, unified diffs, SARIF exports, and per-language fail-closed identifier binding checks guard the edit path.

Trace replay

Bounded neighborhood traces from symbol graph indexes replay context during AI edit workflows and lower the risk of wrong edits.

The next phase is packaging, architecture docs, and real benchmark data.

  • - Split the largest core modules and cut down monolithic reference handling
  • - Publish versioning, wheels, and a standalone server binary
  • - Record large-repo index times, query latency,, and trace replay success rates
Next case study
LokQL
2026 / Local-first virtual query dataflow