MCP Task Orchestrator
Server-enforced workflow discipline for AI agents.
Prompt-based frameworks hope the LLM follows instructions. This one blocks the call if it doesn't.

Task Orchestrator is an MCP server that gives AI coding agents a persistent work item graph with quality gates enforced by the server, not the prompt. It is built for developers running multi-agent or multi-session coding workflows: an orchestrator dispatching sub-agents, a fresh session picking up yesterday's work, or an autonomous loop draining a backlog. It ships as a Docker image, works with any MCP client, and has an optional Claude Code plugin that adds skills and hooks on top.
New here? Start with the illustrated field guide. It explains the ideas on this page in short visual pages, several of them interactive: fire triggers at a phase gate, click a work breakdown through its dependencies, and watch a schema resolve.
The Problem
Multi-agent workflows need infrastructure the model doesn't provide. When an orchestrator dispatches sub-agents across sessions, there's no built-in way to enforce what documentation must exist before work starts, track which agent made which change, or guarantee dependency ordering across a work breakdown. These are structural concerns — they belong in the server, not in prompts.
Task Orchestrator puts them in the server. If a required design note isn't filled, advance_item returns an error naming the missing note. If an upstream dependency isn't complete, the transition is blocked. Every transition and note records who made it. A new session recovers the full state in one call instead of replaying a conversation. And the rules are YAML config, not hardcoded prompts — change them without changing code.
What It Looks Like
Morning — new session, new agent, zero context:
Agent: get_context(since="2025-01-14T17:00:00Z")
→ 2 items in work, 1 blocked, 1 stalled (missing implementation-notes)
→ Recent transitions show orchestrator-1 dispatched 3 sub-agents yesterday
→ Full ancestor chains: "Auth Feature > Login API > Input validation"
Agent: advance_item(transitions=[{ itemId: "a3f2", trigger: "start",
actor: { id: "morning-agent", kind: "subagent", parent: "orchestrator-1" } }])
→ Error: "Gate check failed: required notes not filled for queue phase: requirements"
Agent: manage_notes(operation="upsert", notes=[{ itemId: "a3f2", key: "requirements",
body: "Validate email format, enforce password complexity...",
actor: { id: "morning-agent", kind: "subagent" } }])
→ Upserted. noteProgress: { filled: 1, remaining: 0, total: 1 }
Agent: advance_item(transitions=[{ itemId: "a3f2", trigger: "start",
actor: { id: "morning-agent", kind: "subagent" } }])
→ queue → work. Actor recorded. No context rebuilding.
Quick Start
Prerequisite: Docker installed and running.
1. Register the server
The simplest setup is a per-session STDIO container: no port, no daemon, no REST API. Add it to your project's .mcp.json:
{
"mcpServers": {
"mcp-task-orchestrator": {
"command": "docker",
"args": [
"run", "--rm", "-i",
"-v", "mcp-task-data:/app/data",
"ghcr.io/jpicklyk/task-orchestrator:latest"
]
}
}
}
Claude Code users can register the same shape from the CLI instead:
claude mcp add-json mcp-task-orchestrator '{
"command": "docker",
"args": ["run", "--rm", "-i", "-v", "mcp-task-data:/app/data", "ghcr.io/jpicklyk/task-orchestrator:latest"]
}'
Restart your client. The server creates its SQLite database on first run. Without a config file every tool works in schema-free mode: no gates, no required notes. Add schemas when you want enforcement.
2. Enable workflow schemas
Create .taskorchestrator/config.yaml in your project (see Workflow Enforcement for an example, or use the plugin's /task-orchestrator:manage-schemas skill) and mount that folder into the container:
{
"mcpServers": {
"mcp-task-orchestrator": {
"command": "docker",
"args": [
"run", "--rm", "-i",
"-v", "mcp-task-data:/app/data",
"-v", "/absolute/path/to/your/project/.taskorchestrator:/project/.taskorchestrator:ro",
"-e", "AGENT_CONFIG_DIR=/project",
"ghcr.io/jpicklyk/task-orchestrator:latest"
]
}
}
}
Use an absolute host path. Claude Code's .mcp.json expands only environment variables (${VAR} and ${VAR:-default}), so editor-style placeholders such as ${workspaceFolder} are not substituted. Only the .taskorchestrator/ folder is exposed; the server has no access to the rest of your project.
3. Multi-project setup (HTTP + config-sync)
If you work across several repositories, run one persistent server with the REST API enabled. Each project's .taskorchestrator/config.yaml then syncs into it automatically through the plugin's config-sync hook: no per-project container, no manual mount, and config changes hot-reload without a restart. The plugin's /task-orchestrator:configure-server skill renders this setup interactively; the manual equivalent is:
docker pull ghcr.io/jpicklyk/task-orchestrator:latest
docker run -d --name mcp-task-orchestrator-http --restart unless-stopped \
-v mcp-task-data:/app/data \
-e MCP_TRANSPORT=http -e API_ENABLED=true -e API_AUTH_MODE=none -e API_ALLOW_UNAUTHENTICATED=true \
-p 127.0.0.1:3001:3001 \
ghcr.io/jpicklyk/task-orchestrator:latest
Register it in .mcp.json using the HTTP shape:
{
"mcpServers": {
"mcp-task-orchestrator": {
"type": "http",
"url": "http://localhost:3001/mcp"
}
}
}
And export the client-side variable that tells config-sync where the server is (without it, config-sync silently does nothing):
export TASK_ORCHESTRATOR_API_URL=http://localhost:3001
Alternatively, a client.json (apiUrl only, no token) supplies the URL. Project-mode /task-orchestrator:init writes it beside the project config, only for a loopback server and only as a bare origin with no path, query or fragment (add .taskorchestrator/client.json to .gitignore, the URL is machine-specific); /task-orchestrator:init --user writes the user-level file, which works for any host. The order is the environment variable, then the project-level file, then the user-level file. /task-orchestrator:init sets up a project root for the current directory, and /task-orchestrator:init --user creates a personal root that serves every directory without its own config.
SECURITY: unauthenticated REST means anyone who can reach the port has full read/write/delete
access. This is only safe because the port is published loopback-only (-p 127.0.0.1:3001:3001).
Never publish it on 0.0.0.0 or a wider interface. For shared or networked deployments use bearer
tokens or JWKS auth — see Fleet Deployment.
Prefer to run without Docker? CONTRIBUTING.md covers building the fat JAR from source.
Core Capabilities
The Phase Model
Every work item has a role: queue (not started), work (in progress), review (optional, opt-in per schema), blocked (an upstream dependency is unmet), and terminal (done or cancelled). Each role maps to one or more configurable statuses. Agents move items with advance_item(trigger=...) rather than editing status directly, and that call is where every gate is checked. Schemas attach note requirements to roles: a note declared with role: queue must exist before the item can leave the queue phase.
Illustrated: Phase Gates draws the roles and triggers as a track and includes a gate simulator.
Workflow Enforcement
Schemas define what agents must produce at each phase, and the server blocks progression until it's done. They also set a planning floor: when an agent enters plan mode, the schema tells it what documentation must exist before implementation can start, shaping the plan itself.
# .taskorchestrator/config.yaml
work_item_schemas:
feature-task:
notes:
- key: requirements
role: queue
required: true
description: "Acceptance criteria before starting"
guidance: "Cover: problem statement, acceptance criteria, alternatives considered, test strategy."
skill: "spec-quality"
- key: implementation-notes
role: work
required: true
description: "What was built and why"
With this schema, advance_item(trigger="start") from queue requires requirements to be filled. The server returns an error listing exactly which notes are missing.
The guidance field provides authoring instructions surfaced at the right moment: when an agent is about to fill that note, the response carries the guidance as a guidancePointer. The skill field goes further, naming a skill the agent should invoke before filling the note so the evaluation follows a defined framework rather than freeform prose.
Composable Traits
Traits add cross-cutting note requirements to any schema without duplicating definitions. Define a trait once, apply it to any item type:
traits:
needs-security-review:
notes:
- key: security-assessment
role: review
required: true
description: "Security review of auth, data handling, and access control"
skill: "security-review"
work_item_schemas:
feature-task:
default_traits:
- needs-security-review
notes:
# ... base notes
Every feature-task item automatically inherits the security-assessment requirement. Traits can also be applied per item via the traits parameter on manage_items, so a task touching authentication gets needs-security-review while a CSS cleanup doesn't.
Illustrated: Schemas and Traits shows how a type, its default traits and per-item traits resolve into one gate. Seats covers splitting a phase between several agents.
Persistent Work Item Graph
Everything is a WorkItem in a hierarchical graph. Items nest to any depth and are connected by typed dependency edges. Create an entire work breakdown atomically:
create_work_tree(
root={ "title": "User Authentication" },
children=[
{ "ref": "schema", "title": "Database schema" },
{ "ref": "api", "title": "Login API" },
{ "ref": "tests", "title": "Integration tests" }
],
deps=[
{ "from": "schema", "to": "api" },
{ "from": "api", "to": "tests" }
]
)
When schema reaches terminal, api is automatically unblocked. When all children complete, the parent cascades to terminal. Dependency ordering is enforced by the server, structurally, not by convention.
Illustrated: Work Graph has a clickable version of this tree that shows blocking, unblocking and parent cascades.
Actor Attribution
Every advance_item transition and manage_notes upsert accepts an optional actor claim, and the server records it on the transition or note:
{
"actor": {
"id": "impl-agent-42",
"kind": "subagent",
"parent": "orchestrator-1"
}
}
Query responses include the full delegation chain: which orchestrator dispatched which sub-agent, who wrote which note, who made which transition. Post-mortem debugging becomes a data query rather than conversation archaeology.
To make the claim mandatory, enable actor authentication in config:
actor_authentication:
enabled: true
With the Claude Code plugin installed, a hook then rejects any write that lacks an actor before the call leaves the client. Other MCP clients need to supply the claim by convention. Optional JWKS verification checks that the claim is genuine; see Fleet Deployment.
Session Continuity
No context rebuilding. One call recovers the full picture:
get_context(since="2025-01-15T09:00:00Z", includeAncestors=true)
Returns active items, recent transitions with actor attribution, blocked items, stalled items with their missing notes, and full ancestor chains. A new session has complete state in a single response.
Illustrated: Finding Work maps each question a session has to the read call that answers it, and shows how the next item is ranked.
Notes as Structured Context
Notes are phase-specific documentation attached to work items. An implementation agent reads a concise requirements note scoped to its task instead of scanning broader project context.
Notes are keyed, role-scoped, and queryable:
query_notes(operation="list", itemId="<uuid>", role="work", includeBody=false)
Metadata-only queries (includeBody=false) let agents check what exists without paying the token cost of every note body.
Full-Text Search
Search work items and notes by keyword. Results are relevance-ranked, so an agent looking for related work or picking up after a long gap gets the best matches first.
query_items(operation="search", query="authentication login")
query_notes(operation="search", query="password validation")
Search can be scoped to a subtree and filtered by role, tag, or priority, or run across the whole workspace.
REST API
An optional HTTP layer (API_ENABLED=true) exposes items, notes, dependencies, transitions, config, and real-time SSE events to dashboards, CI systems, and operators. It supports static bearer tokens, JWKS JWT auth, and the unauthenticated loopback mode shown in the Quick Start. It also powers config-sync, so one server can serve many projects with per-project schemas. See the REST API Reference.
Illustrated: Runtime Map covers the server's layers, both deployment shapes, the REST API, claims and actor identity.
Design Philosophy
Task Orchestrator enforces workflow structure without imposing methodology. The server owns the guardrails: role transitions, dependency ordering, gate enforcement, and attribution. Agents own everything else. There are no mandatory planning ceremonies and no opinion on how agents approach implementation. Schemas, traits, and actor authentication are opt-in layers configured in .taskorchestrator/config.yaml. As models gain capabilities, the harness stays out of the way.
Claude Code Plugin
The plugin adds workflow automation on top of the MCP server: skills, hooks, and an orchestration context.
Install:
/plugin marketplace add https://github.com/jpicklyk/task-orchestrator
/plugin install task-orchestrator@task-orchestrator-marketplace
What it adds:
The MCP server works without the plugin. The plugin makes it seamless with Claude Code.
Illustrated: Orchestrating a Run follows one feature from request to finished items. Improvement Loop shows how friction noticed during work becomes a reviewed config change.
Item IDs accept short hex prefixes: itemId="a3f2" works in place of a full UUID once the prefix is unambiguous. See the API Reference for every tool's parameters and response shapes.
Documentation
Technical Stack
- Kotlin with Coroutines
- SQLite + Exposed ORM with FTS5 full-text search, zero configuration
- Flyway versioned schema migrations
- MCP Kotlin SDK with STDIO and Streamable HTTP transports
- Ktor for the REST API and SSE
- Docker for one-command deployment
A layered architecture (Domain → Application → Infrastructure → Interfaces, enforced by a Konsist test against a ratcheting baseline of pre-existing violations) with an extensive JUnit 5 test suite.
License
MIT License — Free for personal and commercial use.