Markdown Vault MCP

Generic markdown vault MCP with hybrid search
Documentation | Config wizard | PyPI | Docker
Give Claude, or any MCP client, a folder of Markdown notes to search, read and write. An Obsidian vault works as it is.
- Hybrid search. Keyword search with SQLite FTS5 and, once an embedding provider is configured, search by meaning, fused by Reciprocal Rank Fusion. Results are short snippets;
read fetches the whole section. See Embeddings.
- Frontmatter as data. YAML frontmatter fields become search filters, and long notes are split at their headings.
- Careful writes. The write tools (
write, edit, append, delete, rename, move_folder, fetch, git_sync, the okf_* tools, create_upload_link) are on by default and hidden when MARKDOWN_VAULT_MCP_READ_ONLY=true. Replacing a file takes the etag from reading it, and per-folder _conventions.md rules reach the client as it writes.
- Git. Optional commit per write with a delayed push, pull or webhook sync, and history and diffs; an overwritten note can be read back at the revision it replaced. See Git integration.
- Links. Backlinks, outlinks, broken links and the path between two notes, for wikilinks and Markdown links alike, with interactive views in clients that render MCP Apps.
- Open Knowledge Format. OKF bundles are recognized, and results carry each note's type, status and trust tier. See OKF.
The tools reference lists every tool. The same engine is a Python library: see the Vault API.
Does it fit?
What the server can reach, what it changes and who gets in is set out in the security model; the block below says who it serves and where it stops.
It suits one person or a small team who keep notes as Markdown files and want Claude to search them, follow their links and write back into them, on their own machine or as a shared server. It reaches the vault folder, plus only what you configure: a git remote, an embedding or summarizing model, and the URLs a fetch call names.
What it assumes:
- One vault per server. Several vaults take one server each, for now (#1232); give each a
MARKDOWN_VAULT_MCP_SERVER_NAME.
- Markdown is what gets searched. Other files, such as PDFs and images, can be read and written as attachments, but their contents are not indexed yet (#1234).
- Embeddings live in memory. Search by meaning holds every vector at 4 bytes × chunks × dimensions: about 70 MB for 23,000 chunks at 768 dimensions, about 900 MiB at ten times that and 1,024 dimensions. Memory, not query time, is the first limit (#1377).
- State sits on local disk. The index and embeddings are files, and the change-tracking file sits beside the index, never inside the vault. Without
MARKDOWN_VAULT_MCP_INDEX_PATH, the index and that file stay in memory and are rebuilt at each start.
Reach for something else for a corpus of hundreds of thousands of chunks (a vector database behind a retrieval pipeline), for mostly scanned or office documents (a document management system with text recognition), or for many users who must not see each other's notes (a multi-tenant knowledge platform): every caller the server admits gets every tool it exposes.
Quick start
Pick the client you use. Each line installs the released version; the Get started tutorials carry on from there.
Claude Desktop. Download the .mcpb bundle from the releases page and open it with Claude Desktop (or Settings › Extensions › Advanced settings › Install Extension…). Claude Desktop asks for the required settings itself. Tutorial.
Claude Code. Two commands inside Claude Code; the second asks for a scope. Tutorial.
/plugin marketplace add pvliesdonk/claude-plugins
/plugin install markdown-vault-mcp@pvliesdonk
A client that runs a command (stdio). To register it in Claude Code, see the Claude Code tutorial.
uv tool install "markdown-vault-mcp"
markdown-vault-mcp serve
A server for remote clients (streamable HTTP). A remote server connects the clients.
docker run --rm -p 8000:8000 --env-file .env ghcr.io/pvliesdonk/markdown-vault-mcp:latest
A compose.yml ships at the repository root and runs as-is: copy .env.example to .env, then docker compose up -d. Deploy covers authentication, OIDC, a reverse proxy and system packages (.deb/.rpm on the releases page). The server answers /health and /health/ready outside the MCP mount, and its get_server_info tool reports the running version.
The plain package covers keyword search, the write tools and git. Search by meaning, the file watcher and the summarize tool need extras; [all] installs every one:
uv tool install "markdown-vault-mcp[all]"
The Docker image, the .mcpb bundle and the Claude Code plugin already include them. Installation lists each extra.
Configuration
Everything is configured through environment variables with the MARKDOWN_VAULT_MCP_ prefix. The ones most installs set:
Every variable the server reads, the shared ones included, is in the configuration reference; .env.example lists the same surface in copy-paste form, and the config wizard writes one for your deployment.
Documentation
- Security model: what the server can reach, what it changes and who gets in.
- Get started: a first success with your client.
- Deploy: Docker, authentication, OIDC, reverse proxy.
- Use: the features, for real tasks.
- Reference: configuration, tools, resources, prompts, command line.
- Upgrade: release channels, and what an upgrade changes for your clients and your data.
- Contribute: local development, secrets, where a fix belongs;
CONTRIBUTING.md and SECURITY.md at the root.
Design decisions
- Document identity is the relative path with
.md extension; frontmatter is optional by default (REQUIRED_FIELDS opts into enforcement).
- Hybrid search uses Reciprocal Rank Fusion over the FTS5 and vector result lists, with diversity-aware ranking capping chunks per document.
- Tool semantics mirror Claude Code's Read/Write/Edit patterns, so LLM clients drive the vault with habits they already have.
- The library is synchronous; the MCP layer wraps calls in
asyncio.to_thread(). File writes return after saving, while index updates run in the background. Index-dependent mutations wait for prior writes; see index freshness.
- Indexing is hash-based: unchanged files are never re-parsed, and any change to how stored rows derive from a note's bytes bumps
INDEX_SEMANTICS_VERSION so deployed vaults rebuild themselves once on upgrade.
The full decision log lives in the design document.
Links