
data-gov-il-mcp
Production-grade Model Context Protocol (MCP) server for Israeli Government Open Data from data.gov.il.
This server gives MCP-compatible clients structured access to Israeli government datasets, resources, tags, organizations, and tabular records through the official CKAN API. It is written in TypeScript, validates inputs and outputs with Zod, returns JSON-first tool responses, and supports both local stdio and remote Streamable HTTP transports.
Highlights
Quick Start
Claude Desktop / Local Stdio
Use the published package directly:
{
"mcpServers": {
"data-gov-il": {
"command": "npx",
"args": ["-y", "data-gov-il-mcp"]
}
}
}
Streamable HTTP
npm install -g data-gov-il-mcp
data-gov-il-mcp-http
Then configure your MCP client:
{
"mcpServers": {
"data-gov-il": {
"url": "http://localhost:3664/mcp"
}
}
}
Docker
docker run -p 3664:3664 ghcr.io/davidosherproceed/data-gov-il-mcp:latest
For local development:
npm install
npm run build
npm run start:stdio
# or
npm run start:http
Catalog Discovery Layer
The server ships with a committed catalog snapshot at src/data/catalog/catalog.snapshot.json. The snapshot is generated from data.gov.il and bundled into the build. At startup, it is validated and indexed in memory.
The discovery layer powers:
find_datasets catalog-first search with live CKAN fallback.
list_all_datasets instant local enumeration with optional organization filter.
list_organizations instant local organization list with dataset counts.
list_available_tags and search_tags from real CKAN tag/facet data.
- Dynamic completions for
datagov://dataset/{id}.
datagov://tags and datagov://catalog/stats resources.
It includes:
- Hebrew normalization, including nikud removal and final-letter normalization.
- Tokenization and trigram indexes for fuzzy matches and typo tolerance.
- Dataset, tag, and organization maps.
- Tag-to-dataset and organization-to-dataset indexes.
- Tag co-occurrence for related tag suggestions.
- Weighted ranking across exact, token, tag, organization, and fuzzy signals.
Refresh the snapshot:
npm run catalog:refresh
npm run build
A scheduled GitHub Actions workflow (.github/workflows/catalog-refresh.yml) refreshes the snapshot and opens a PR when catalog data changes.
More detail: docs/catalog-discovery-layer.md.
All successful tool responses return:
structuredContent: the typed JSON object.
content[0].text: the same object serialized as JSON for text-only clients.
Elicitation in find_datasets
find_datasets can optionally expose an interactive parameter:
{
"query": "תחבורה",
"interactive": true
}
This is only registered when:
MCP_ENABLE_ELICITATION=true
When enabled, compatible clients may show a clarification form for broad searches. For example, the server can ask the user to narrow many matching datasets by publisher organization. If the client does not support Elicitation, the user declines, or the request times out, the tool falls back to normal search results.
This is disabled by default because MCP client support varies.
Resources
The server also implements resource subscriptions in a minimal standards-compliant way:
- Advertises
resources.subscribe.
- Handles
resources/subscribe and resources/unsubscribe.
- Sends
notifications/resources/updated only for resources a client subscribed to.
- Does not poll CKAN in real time.
Prompts
Prompt arguments use MCP completions. Domain focus suggestions are curated, and organization completions are catalog-backed.
Optional MCP Client Features
These features are off by default. Enable them only when your target MCP client supports them and you want the server to expose them.
Client support differs:
- Cursor supports Elicitation, but does not currently expose Sampling.
- Claude Code supports Elicitation in recent versions.
- Claude Desktop supports many MCP features, but Elicitation support is not reliable/available.
- Sampling availability varies; the server always falls back safely.
Configuration
Copy .env.example to .env:
cp .env.example .env
Core
CKAN
HTTP Hardening
Authentication
Service Identity
Recommended Workflows
Find and Query a Dataset
- Use
find_datasets with natural Hebrew or English terms.
- Use
get_dataset_info or list_resources for a chosen dataset.
- Pick a resource with
datastore_active=true.
- Use
search_records with limit=5 first to inspect fields.
- Add
filters, fields, sort, distinct, or pagination as needed.
Example flow:
find_datasets({ "query": "מחיר למשתכן" })
get_dataset_info({ "dataset": "mechir-lamishtaken" })
search_records({
"resource_id": "7c8255d0-49ef-49db-8904-4cf917586031",
"limit": 5,
"include_total": true
})
search_tags({ "keyword": "דיור", "limit": 5 })
find_datasets({ "query": "תחבורה", "tags": "תחבורה ציבורית" })
Use Interactive Discovery
Requires:
MCP_ENABLE_ELICITATION=true
Then an agent may call:
find_datasets({ "query": "תחבורה", "interactive": true })
Compatible clients may show a form asking the user to narrow results.
Use Client-Side Summaries
Requires:
MCP_ENABLE_SAMPLING=true
Then:
summarize_dataset({ "dataset": "mechir-lamishtaken", "language": "he" })
If Sampling is unavailable, the tool returns the dataset metadata and sampling.used=false.
Development
npm install
# Type-check
npm run typecheck
# Lint
npm run lint
# Test
npm test
# Build
npm run build
# Refresh local catalog snapshot
npm run catalog:refresh
Run locally:
# stdio
npm run build
npm run start:stdio
# HTTP
npm run build
npm run start:http
Enable optional features locally:
MCP_ENABLE_ELICITATION=true MCP_ENABLE_SAMPLING=true npm run start:http
On PowerShell:
$env:MCP_ENABLE_ELICITATION="true"
$env:MCP_ENABLE_SAMPLING="true"
npm run start:http
Project Structure
src/
auth/ Authentication providers and Express middleware
bin/ stdio and HTTP entry points
cache/ In-memory TTL/LRU cache
catalog/ Snapshot validation, indexing, fuzzy search, CatalogService
ckan/ Typed CKAN API client and CKAN response types
config/ Zod env config, constants, server identity
core/ Dependency container, MCP server factory, lifecycle
data/catalog/ Committed catalog snapshot artifact
formatting/ JSON response builders and guidance text
observability/ Pino logger
prompts/ MCP prompt definitions, templates, registration
resources/ MCP resources, templates, subscriptions
services/ Domain services for CKAN data access
tools/ MCP tool definitions and Zod schemas
transports/ stdio and Streamable HTTP transports
tests/
fixtures/ Test fixtures
unit/ Unit tests
scripts/
refresh-catalog.ts
docs/
catalog-discovery-layer.md
MIGRATION.md
Docker
docker build -t data-gov-il-mcp .
docker run -p 3664:3664 data-gov-il-mcp
With optional features:
docker run -p 3664:3664 \
-e MCP_ENABLE_ELICITATION=true \
-e MCP_ENABLE_SAMPLING=true \
data-gov-il-mcp
Quality
The project is expected to pass:
npm run typecheck
npm run lint
npm test
npm run build
Current implementation includes unit coverage for environment parsing, auth providers, CKAN errors, cache, formatting, catalog text normalization/fuzzy/index/search logic, snapshot validation, services, resources, subscriptions, and HTTP Host/Origin guard.
License
MIT