MinerU Document Explorer

by opendatalab

Agent-native knowledge engine — search, deep-read, and build knowledge bases from Markdown, PDF, DOCX, and PPTX. Supports hybrid search with BM25, vector search, and LLM reranking. Optional external data files include LLM models auto-downloaded on first use (~300MB to ~1.1GB). Python packages are optionally required for PDF, DOCX, and PPTX processing.

Education & sciencestdio or Streamable HTTPCommunity

Repository-wide counts · Cached 2026-04-07

Overview

The MinerU Document Explorer MCP server is a publicly available project. Review the upstream repository for installation instructions, supported tools, compatibility, permissions, and current maintenance status.

Configuration

Configuration, transport, authentication, and runtime requirements vary by project. Open the repository before connecting and use the smallest set of credentials and permissions required.

Open the MinerU Document Explorer repository to read the latest documentation.

KEEP EXPLORING

Compare source, connection, and authentication details before choosing an implementation.

View the complete category

Deep Research

u14app

Community

Deep Research 使用强大的 AI 模型快速生成深入的研究报告。支持 SSE API 和 MCP 服务器。需要在 .env 文件中配置环境变量以设置服务器端的 API 密钥和相关参数。

TorchLeet

Exorust

Community

TorchLeet provides 68 PyTorch problems from real ML/AI interviews at companies like Google, Meta, and Anthropic. It includes an AI Tutor MCP server that gives AI assistants access to problems, hints, prep plans, and learning paths with a no-spoilers teaching style.

Zotero MCP

54yyyu

Community

用于 Zotero 的模型上下文协议(MCP)服务器,将您的 Zotero 研究库与 Claude 及其他 AI 助手连接。支持本地和 Web API 访问、PDF 注释提取以及高级搜索功能。完整本地 API 功能需要 Python 3.10 及 Zotero 7 以上版本。配置可以通过环境变量或 JSON 配置文件进行设置。

mcp-brasil

mcp-brasil

Community

MCP Server for 70 Brazilian public data sources covering economy, legislation, transparency, judiciary, elections, environment, health, education, public security, and more. Some APIs require optional API keys configured via environment variables (e.g., TRANSPARENCIA_API_KEY, DATAJUD_API_KEY, META_ACCESS_TOKEN).

FROM THE SOURCE

Repository README

Build-time snapshot · Retrieved 2026-10-05

View original

logo MinerU Document Explorer

Agent-native knowledge engine — search, deep-read, and build knowledge bases
from Markdown, PDF, DOCX, and PPTX.

npm license CI stars

中文文档 · MCP Setup · CLI Reference · Demo · Contributing


🤔 Why MinerU Document Explorer?

MinerU Document Explorer equips your agent with three tool suites — Retrieve, Deep Read, and Ingest — closing the full knowledge loop:

Overview of MinerU Document Explorer

  • 🔍 Retrieve — Cross-collection search: BM25, vector, and hybrid with LLM reranking and query expansion
  • 📖 Deep Read — Navigate inside a single document without loading the whole file: table of contents, section reading, inline search, and element extraction
  • 📝 Ingest — Build and maintain a LLM wiki from raw documents, following the Karpathy LLM Wiki pattern

Developed by the MinerU team, building on QMD and Karpathy's LLM Wiki.

💡 What can you do with it?

  • Research assistant — Point it at a folder of PDFs and ask your AI agent to survey the literature, compare methods, or find specific results across papers
  • Project knowledge base — Index your codebase, docs, and design specs; let the agent answer questions about architecture, find relevant code, or trace requirements
  • Study companion — Upload textbooks or lecture notes; use doc_toc + doc_read to navigate, doc_grep to search, and build a wiki of key concepts
  • Enterprise document search — Index contracts, reports, or manuals and search them with natural language queries enhanced by LLM reranking

See it in action: The demo/ folder contains a complete end-to-end example — an AI agent automatically reads ~10 arXiv papers on RAG, builds an interlinked wiki knowledge base, and writes a research survey. See the Demo Guide for step-by-step instructions.

🚀 Quick Start

Agent-Assisted Setup: If you're using an AI agent (Claude Code, Cursor, etc.), simply ask it to help you deploy MinerU Document Explorer and install skills — the agent can handle the entire setup process for you, including MCP configuration.

Follow the https://github.com/opendatalab/MinerU-Document-Explorer/blob/main/docs/quickstart.md to install MinerU Document Explorer and walk the user through configuration.

📖 Document Deep Reading

Navigate and search within a single document without reading the whole file:

# View document structure
qmd doc-toc papers/attention-is-all-you-need.pdf

# Read specific sections by address
qmd doc-read papers/attention-is-all-you-need.pdf "line:45-120"

# Search within one document
qmd doc-grep papers/attention-is-all-you-need.pdf "self-attention"

🔌 MCP Server — 15 Tools for AI Agents

Integrate with AI agents via Model Context Protocol.

MCP Server vs CLI: The MCP server runs as a persistent process — LLM models (embeddings, reranker, query expansion) are loaded once and stay in memory across requests. CLI commands like qmd query must reload all models on every invocation, adding ~5–15 s of startup overhead each time. For agent workflows, always prefer the MCP server.

Two transport modes:

Mode Command Best for
stdio qmd mcp Claude Desktop, Claude Code — client spawns and manages the process
HTTP daemon qmd mcp --http --daemon Cursor, Windsurf, VS Code, multi-client setups — one shared persistent server
# Start the HTTP daemon (recommended — models stay loaded across all requests)
qmd mcp --http --daemon             # default port 8181
qmd mcp --http --daemon --port 8080 # custom port

# Verify server is running
curl http://localhost:8181/health

# Stop the daemon
qmd mcp stop

Client Configuration

Cursor — add to .cursor/mcp.json (project) or ~/.cursor/mcp.json (global)

Option A — stdio (Cursor manages the process lifecycle):

{
  "mcpServers": {
    "qmd": {
      "command": "qmd",
      "args": ["mcp"]
    }
  }
}

Option B — HTTP (run qmd mcp --http --daemon first; models stay loaded, faster responses):

{
  "mcpServers": {
    "qmd": {
      "url": "http://localhost:8181/mcp"
    }
  }
}
Claude Desktop — add to ~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "qmd": {
      "command": "qmd",
      "args": ["mcp"]
    }
  }
}
Claude Code — add to ~/.claude/settings.json or run claude mcp add qmd -- qmd mcp
{
  "mcpServers": {
    "qmd": {
      "command": "qmd",
      "args": ["mcp"]
    }
  }
}
Windsurf / VS Code / Other MCP Clients

For stdio transport, use "command": "qmd", "args": ["mcp"] in your client's MCP configuration.

For HTTP transport, start qmd mcp --http --daemon and point your client to http://localhost:8181/mcp.

See MCP setup guide for all 15 tools and HTTP transport details.

Agent Skills

MinerU Document Explorer ships with a built-in Agent Skill that teaches AI agents how to use the full tool suite effectively — decision trees, usage patterns, and best practices for all 15 MCP tools.

# Install the skill (works with both npm and source installs)
qmd skill install              # local project (.agents/skills/)
qmd skill install --global     # global (~/.agents/skills/)

# Or from source repo
claude skill add ./skills/mineru-document-explorer/SKILL.md

📊 How It Compares

MinerU Doc Explorer LlamaIndex Obsidian NotebookLM
Runs 100% locally ✅ ⚠️ LLM APIs ✅ ❌ Cloud
Agent integration (MCP) 15 tools Plugin ❌ ❌
Deep reading within docs ✅ ❌ ❌ ✅
Wiki knowledge compilation ✅ ❌ Manual ❌
Formats MD, PDF, DOCX, PPTX Many MD PDF, URL
Search pipeline BM25 + vec + rerank Configurable Basic Proprietary
Zero-config search ✅ qmd search ❌ Plugin N/A
Open source MIT MIT Partial ❌

⚙️ Requirements

Requirement Notes
Node.js >= 22 or Bun Runtime
Python >= 3.10 Document processing (pymupdf, python-docx, python-pptx)
macOS brew install sqlite for extension support

📄 Document Processing Setup

Python 3.10+ is required for document processing (PDF, DOCX, PPTX):

# Check Python version
python3 --version  # needs >= 3.10

# Install required Python packages
pip install pymupdf python-docx python-pptx

# Verify
python3 -c "import pymupdf; import docx; import pptx; print('OK')"
MinerU Cloud — high-quality PDF extraction for scanned documents and complex layouts (optional)
pip install mineru-open-sdk
export MINERU_API_KEY="your-key"  # get from https://mineru.net

When MINERU_API_KEY is set, MinerU Cloud is automatically used as the primary PDF provider with PyMuPDF as fallback.

For advanced configuration (custom providers, local VLM models, GPT PageIndex), create ~/.config/qmd/doc-reading.json:

{
  "docReading": {
    "providers": {
      "fullText": { "pdf": ["mineru_cloud", "pymupdf"] }
    },
    "credentials": {
      "mineru": { "api_key": "your-api-key" }
    }
  }
}

🤖 LLM Models (auto-downloaded on first use)

Model Purpose Size
embeddinggemma-300M Vector embeddings ~300 MB
qwen3-reranker-0.6b Re-ranking ~640 MB
qmd-query-expansion-1.7B Query expansion ~1.1 GB

Models are only needed for qmd embed, qmd vsearch, and qmd query. qmd search runs BM25 retrieval.

📚 Documentation

🎯 Demo Guide End-to-end example: agent-driven RAG research survey
📖 CLI Reference All commands, options, output formats
🔌 MCP Server Setup, 15 tools, HTTP transport
📦 SDK / Library TypeScript API, types, examples
🏗️ Architecture Search pipeline, scoring, data schema, chunking
🤝 Contributing Development setup, code style, how to contribute

❤️ Acknowledgments

MinerU Document Explorer builds upon these foundational projects:


📝 Changelog

v1 — 2026-04-07 (Current)

Rebuilt from an OpenClaw agent skill into a full agent-native knowledge engine: npm package (npm install -g mineru-document-explorer), qmd CLI, MCP server with 15 tools across three groups (Retrieval / Deep Reading / Knowledge Ingestion), multi-format support (MD, PDF, DOCX, PPTX), hybrid search (BM25 + vector + LLM reranking), and LLM Wiki knowledge base pattern.

v0 — 2026-03-30 (Previous)

OpenClaw-native agent skill (doc-search CLI). Four capabilities: Logic Retrieval, Semantic Retrieval, Keyword Retrieval, Evidence Extraction. See the v0 repository.