Back to Discover

windbg-mcp

connector

glslang

WinDbg/DbgEng over MCP: crash dumps, live user & kernel, driver IOCTLs, and TTD.

View on GitHub
0 starsSynced Aug 3, 2026

Install to Claude Code

/plugin marketplace add glslang/windbg-mcp

README

windbg-mcp

CI License: MIT CodeRabbit Pull Request Reviews Latest release Platform: Windows x64

An MCP server that exposes WinDbg/DbgEng to AI agents (Claude Code, Claude Desktop, Cursor, …) over stdio. It drives a live debugger engine for user-mode, kernel-mode, crash-dump, and Time Travel Debugging (TTD) workflows.

The low-level engine bindings live in win-kexp (src/dbgeng.rs); this crate adds process-per-session engine supervision and the rmcp tool surface on top.

Architecture

One debug session per process. dbgeng.dll holds a single debuggee session per process — that is why .opendump replaces the target rather than opening a second one — so this server runs the MCP protocol in a supervisor process and each open target in its own engine worker child process. Two things follow, and they are why it is built this way:

  • A session that cannot be unwound costs a process, not the server. A live-kernel attach waits for its target with WaitForEvent(INFINITE), and nothing can interrupt a wait that has not yet connected — so a guest that never dials in blocks forever. That blocks one worker, which end_session can terminate. It used to block the server's only engine thread, and the only recovery was restarting the server.

  • Sessions are concurrent. Triage a crash dump while a kernel attach is live; keep a TTD trace open while you look at another. Up to four at once.

  • engine.rs — the supervisor: the session registry, worker spawn/teardown, and the routing that turns a session_id into "which worker". Each session has one queue with one consumer, so calls against a session are serialized — one runs at a time, and the one running finishes before the next starts. Serialized is not ordered: two calls submitted before either has answered reach that queue in whichever order wins the race, so await each result before sending the next call that depends on it.

  • worker.rs — the child process. The DebugEngine is created on, and confined to, one OS thread inside it (DbgEng requires serialized, single-thread access, and WaitForEvent must run on the session-owning thread). A catch_unwind guard turns a panic in one operation into a failed call rather than a dead session.

  • proto.rs — the line-delimited JSON protocol between the two. A closure cannot cross a process boundary, so what used to be closures marshalled onto the engine thread are now serializable operations — deliberately tool-shaped rather than DbgEng-shaped, so a tool that is several engine calls (reachable_from_dispatch's call-graph walk) stays one indivisible job.

  • server.rs — the MCP tools (see below), built with rmcp's #[tool_router]/#[tool_handler].

  • ttd.rs — locates TTD.exe and launches trace recording.

  • main.rs — role selection (supervisor or worker), tokio + stdio transport. Logs go to stderr (stdout is the JSON-RPC channel); workers inherit the supervisor's stderr, so everything lands in the same place. Workers never outlive the connection: a disconnect asks every session to release its target — all of them concurrently — waits five seconds, and terminates only the workers that have not finished by then; a worker also exits on its own if its stdin closes. Which of those two endings a session gets matters for a live kernel. DbgEng leaves a detached-but-halted kernel frozen, so a worker that releases its target leaves the machine running, while a worker that is terminated leaves it stopped. Five seconds is enough for an idle session and for most busy ones, but a session in the middle of long work may not make it, so end a live kernel session with end_session — which allows considerably longer — rather than relying on the disconnect.

MCP protocol revision: built on rmcp 3.x, this server accepts every revision that SDK knows — 2026-07-28 and the initialize-handshake ("legacy") era before it (2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05) — and serves whichever the client selects. A 2026-07-28 client gets the stateless, per-request model (server/discover, resultType, per-request _meta) and may open with server/discover instead of initialize; older clients keep the handshake, and a client that offers an unknown revision is answered with 2025-11-25.

Requirements

  • Windows x64 (host bitness must match the target).
  • dbgeng.dll / dbghelp.dll — present in System32 on modern Windows 11 (verified with 10.0.26100). This is enough for live user-mode/kernel debugging and crash-dump analysis.
  • For crash-dump !analyze (and any other !-extension command), the engine needs the WinDbg winext\ extensions bundled next to the binary — System32's engine ships none, so !analyze would return "No export analyze found". See Bundling the WinDbg engine below.
  • For Time Travel Debugging (.run) replay, the System32 engine is not enough — it rejects .run traces (0x80070057). You need the WinDbg engine (which bundles the TTD replay components) loaded next to the binary — see TTD engine below.
  • TTD.exe (the standalone Time Travel Debugging recorder) for record_trace — ships with the WinDbg / TTD store packages; put it on PATH.
  • A reachable symbol server (e.g. srv*https://msdl.microsoft.com/download/symbols) for symbol-name queries like ttd_calls("ucrtbase!_stdio_common_vfprintf"). Offline, address-based queries and the data model still work; symbol names won't resolve.
  • Administrator for live kernel debugging and TTD recording (not for replay).

Build or download

Prebuilt Windows x64 binaries are attached to each GitHub release as windbg-mcp-vX.Y.Z-windows-x64.zip (with a SHA256SUMS.txt to verify the download against — the skill's setup.md snippet does this for you) — no Rust toolchain needed. To build from source instead:

cargo build --release

win-kexp is fetched automatically as a git dependency from glslang/win-kexp — no sibling checkout needed.

cargo test covers the unit tests plus an end-to-end smoke test that drives the built binary over stdio. Run it after a dependency bump or an MCP spec revision — see docs/smoke-test.md.

Bundling the WinDbg engine

Needed for two things: TTD .run replay (System32's engine rejects traces with 0x80070057) and crash-dump !analyze (which lives in the winext\ extensions that System32 doesn't ship). DebugCreate binds to whichever dbgeng.dll the loader finds first, and the app directory is searched before System32, so the copied WinDbg engine (which replays TTD traces and ships the extensions) wins. One-time, from the installed WinDbg store package:

$wd  = (Get-AppxPackage Microsoft.WinDbg).InstallLocation + "\amd64"
$dst = "C:\workspace\windbg-mcp\target\release"
Copy-Item "$wd\dbgeng.dll","$wd\dbghelp.dll","$wd\dbgcore.dll","$wd\dbgmodel.dll",`
          "$wd\symsrv.dll","$wd\msdia140.dll" $dst -Force
Copy-Item "$wd\ttd"    "$dst\ttd"    -Recurse -Force   # TTDReplay*.dll, TtdExt.dll, TTDAnalyze.dll, ...
Copy-Item "$wd\winext" "$dst\winext" -Recurse -Force   # ext.dll (!analyze), kext.dll, … — crash-dump triage
  • The ttd\ subdir provides the @$cursession.TTD / @$curprocess.TTD data model and the !tt time-travel commands.
  • The winext\ subdir provides ext.dll (which exports !analyze) and the other !-extensions. open_dump runs .load ext for you, but note the unqualified !analyze does not resolve on this minimal engine — use the module-qualified !ext.analyze -v for crash-dump triage. Without winext\, !analyze returns "No export analyze found".
  • msdia140.dll is required for PDB symbols. Without it, dbghelp can't parse any PDB (dia error 0x8007007e) and silently falls back to export symbols — which makes module!name lookups (and so ttd_calls("ucrtbase!__stdio_common_vfprintf")) fail even with the right PDB in the cache. symsrv.dll is needed to read a symbol-store cache.

(cargo clean wipes target\, so re-copy after one.) Live and dump debugging work with or without the TTD engine; PDB symbol names need msdia140.dll + a symbol path (execute.sympath srv*C:\ProgramData\Dbg\sym*https://msdl.microsoft.com/download/symbols).

Use with an MCP client

Point your client at the built binary, e.g. Claude Code:

// .mcp.json  (or claude_desktop_config.json under "mcpServers")
{
  "mcpServers": {
    "windbg": {
      "command": "C:\\workspace\\windbg-mcp\\target\\release\\windbg-mcp.exe"
    }
  }
}

As a Claude Code plugin

This repo is also a single-plugin Claude Code marketplace: installing it registers the windbg MCP server and a windbg-debugging skill that knows how to drive it (setup, crash-dump, live/kernel, and TTD playbooks).

/plugin marketplace add glslang/windbg-mcp
/plugin install windbg-mcp@windbg-mcp

The plugin ships source, not a binary, so after installing you still put the server binary in place — download a prebuilt release or build from source — and (for .run replay and crash-dump !analyze) bundle the WinDbg engine — the skill's setup.md walks through it, and it mirrors the Build or download and Bundling the WinDbg engine sections above. Then /reload-plugins to connect the server. The plugin points at ${CLAUDE_PLUGIN_ROOT}/target/release/windbg-mcp.exe.

From the official MCP registry

The server is listed in the official MCP registry as io.github.glslang/windbg-mcp. Clients that support the registry — or that install MCPB bundles directly — can add it by name: the client downloads that release's .mcpb bundle, verifies its SHA-256, and wires up the windbg-mcp.exe inside it as an stdio server, with no Rust build or manual binary placement.

The bundle is Windows x64 only and ships just the server binary, so the one-time engine setup still applies — for TTD .run replay and crash-dump !analyze, drop the WinDbg engine DLLs next to the client-extracted windbg-mcp.exe (the skill's setup.md covers it). Basic live and crash-dump work runs on the in-box System32 engine without them.

Releasing

The plugin sets an explicit version in .claude-plugin/plugin.json, so users only receive an update when that version changes — pushing commits alone does not trigger one. To cut a release, bump version in plugin.json and Cargo.toml, bump the release badge near the top of this README, add a matching entry to CHANGELOG.md, and tag the commit vX.Y.Z. Run claude plugin validate . --strict before publishing. Pushing the tag runs release.yml, which verifies the tag matches both manifest versions and the README badge, builds windbg-mcp.exe, and attaches the zip + SHA256 checksum to the GitHub release. It also builds an MCPB bundle (windbg-mcp-vX.Y.Z-windows-x64.mcpb, described by packaging/mcpb/manifest.json) and publishes a server.json entry to the official MCP Registry (io.github.glslang/windbg-mcp) with the mcp-publisher CLI over GitHub OIDC — no secrets. CI stamps the release version into both files and the bundle's SHA-256 into server.json, so neither is part of the manual bump list above. The zip also gets a signed build-provenance attestation tying it to the workflow run that built it — verify with:

gh attestation verify <zip> --repo glslang/windbg-mcp `
   --signer-workflow glslang/windbg-mcp/.github/workflows/release.yml

(--repo alone only proves the attestation came from some workflow in this repo; --signer-workflow pins it to the release workflow.)

Walkthroughs

  • docs/crash-dump-walkthrough.md — triaging a real kernel minidump (docs/samples/052126-34312-01.dmp): a 0x9F DRIVER_POWER_STATE_FAILURE traced to nvlddmkm.sys via !ext.analyze -v and a manual device-stack walk, with the real outputs and the partial-minidump (0x80040205) gotcha.
  • docs/ttd-walkthrough.md — a hands-on tour of the TTD tools against the xusheng6/TTD_lab helloworld sample: opening a .run, surveying events/threads, forward/reverse navigation, memory analysis, and counting printf calls with symbols (with the real outputs and the gotchas). It maps each tool to the lab's exercises.
  • docs/flareauthenticator-ttd-walkthrough.md — a full TTD → Z3 solve of an obfuscated Qt crackme (Flare-On 12 #8). Defeats an anti-analysis env guard, records a wrong-guess run, and uses ttd_calls/ttd_memory/reverse-navigation to peel control-flow flattening + computed calls + encrypted strings down to the exact check: a per-keystroke rolling hash that reduces to a pure weighted sum. The 25 weights come from the replay, the 250 g values from debugger function-evaluation, and Z3 finds a satisfying code — which reveals the flag (the code is intentionally non-unique). Runnable solver in examples/flareauthenticator/; recorded terminal session in docs/flareauthenticator.cast (asciinema play) — rendered as a GIF in the walkthrough.
  • docs/driver-ioctl-walkthrough.md — enumerating a driver's IOCTL surface and deciding user-mode reachability on a live KDNET kernel: driver_object/uf to recover the \Driver\mountmgr dispatch switch, decode_ioctl for the access tiers, the device DACL parsed from memory, and an ioctl_trace sweep — ending with a reachability report (which codes a standard user can reach vs. what the I/O manager blocks).

Tools

GroupTools
Sessionopen_dump, open_trace, attach_kernel_local, attach_kernel, attach_process, launch, end_session, session_status
Stateregisters, read_memory, backtrace, modules, threads, disassemble, dx
Controlgo, step_over, step_into, set_breakpoint
TTD navstep_back (t-), step_over_back (p-), reverse_go (g-), goto_position (!tt)
TTD analysisttd_calls, ttd_memory, ttd_events, index_trace, record_trace
Driver IOCTLdecode_ioctl, driver_object, device_object, irp_stack, ioctl_trace, reachable_from_dispatch
Rawexecute — run any debugger command, returns full text output

Sessions and session handles

The six Session tools that open a target (open_dump, open_trace, attach_kernel_local, attach_kernel, attach_process, launch) each create a session — one engine worker process holding one target — and return a session_id. Every tool that touches a target accepts that id as an optional argument, and it is what routes the call to the right worker.

Sessions are independent. Opening a second target does not disturb the first, a call against one does not queue behind work in another, and ending one leaves the rest alone. Up to 4 at once; at the limit a new open reclaims the oldest idle session, and if every session has a call in flight the open is refused with the list rather than picking a victim. Sessions end when you end_session them or when the client disconnects — a disconnect is treated as end_session on everything, so no debugger process is left behind. It gives each session a shorter grace than end_session does, though, so end a live kernel session explicitly if you can: one still busy at disconnect is terminated, and a terminated kernel session leaves its target halted.

Omit session_id and a call goes to the current session: the most recently opened one that will still accept work. That is the pre-handle behaviour and it still holds. What supplying the id buys is that the call can never land on a target you did not open — it fails loudly instead. decode_ioctl (pure) and record_trace (independent of any debug session) do not take one.

session_status lists every session — what it is, what state it is in, how long it has been there, and which one is current — or reports on one you name. It never queues on any worker, so it answers even while a session is parked.

Recovering a session that is stuck. A per-call timeout abandons the wait, not the job, so a call that reports a timeout may still be running. The case that matters is attach_kernel: it waits for the target to dial in with no timeout, and DbgEng cannot interrupt a wait that has not yet connected — so a guest that is powered off, not booted with debugging enabled, or pointed at the wrong host/port/key never arrives, and that wait never ends. session_status distinguishes a link that is still coming up (normal, ~25s for a KDNET resync) from one that has been waiting far longer than a healthy attach ever takes. For the second, end_session is the recovery: it asks the worker to let go, and terminates the worker process if it will not. Do not re-run the open while it is still waiting — the target was already claimed, so that would attach a second time.

Two caveats, both in the command hatches, and both now confined to a single session. The typed tools announce their own transitions, but execute can replace its session's target directly (.opendump, .attach, .detach, .kill, .restart, .abandon, .remote, q/qd/qq), and those commands retire that session's handle: calls passing it are refused, while calls that pass no id still reach the worker. dx is the second hatch — the data model reaches Debugger.Utility.Control.ExecuteCommand, which runs any command, so an expression that touches command execution retires the handle too, conservatively, since the command it runs is a runtime string this server never sees.

Both matches are deliberately biased toward retiring: over-matching costs one re-open, under-matching would let a stale handle through. Neither can be exhaustive — DbgEng has more ways to reach the target than a name list can enumerate, and the data model is extensible — so inside execute and dx a handle is a strong hint rather than a guarantee. Everywhere else it is a guarantee, and it is enforced at the front of that session's queue, after everything queued ahead of it: checking on the caller's side would leave a window in which an execute { ".opendump …" } already queued ahead retires the handle between a caller's check and its call.

Typed operands are operands, not commands

The typed tools build debugger commands by interpolation (u {address}, bp {expression}, !drvobj {name} 7), so those parameters refuse ;, line breaks, and " — the last everywhere except dx, whose data-model expressions use quoted literals legitimately.

Two things go wrong without that. DbgEng treats ; as a command separator, so disassemble { address: "rip; .opendump C:\other.dmp" } would replace the debug target from a tool that reports itself read-only. And bp <location> "command" is real WinDbg syntax — ioctl_trace builds exactly that form — so a quote in a breakpoint location arms a command that runs on every hit, replacing the target at some arbitrary later moment, outside any tool call and outside anything that could retire the session handle.

Nothing legitimate is lost: these parameters were always single operands. Use execute to run a command list — it is annotated destructive and retires the handle when a command changes the target.

Error reporting

Anything scoped to a session comes back as a normal tool result with isError: true and the text intact, so the model can read it and correct itself: a failed debugger operation (an unresolvable symbol, an unreadable address, a target that never stopped), a refused handle, a timeout, a session whose worker is gone. Each has a next move. The only JSON-RPC protocol error is the one failure that is the server's rather than a session's — no engine worker could be started at all.

The forward (go/step_over/step_into) and reverse (reverse_go/step_over_back/step_back) control tools mirror a debugger UI's F9/F8/F7 and Shift+F9/F8/F7, so an agent can drive a trace in both directions and jump anywhere with goto_position. All of these issue the command and pump the engine to the next stop (a plain Execute only sets the run state — it doesn't move the target), which is what makes both live stepping and TTD forward/reverse navigation actually advance.

ttd_calls/ttd_memory/ttd_events are convenience wrappers over the TTD data model: ttd_calls and ttd_memory query @$cursession.TTD.{Calls,Memory} (every call to a function / every access to an address range), and ttd_events queries @$curprocess.TTD.Events (the module/thread/exception timeline). For anything else, dx evaluates arbitrary data-model/LINQ expressions, e.g. @$cursession.TTD.Calls("ntdll!NtCreateFile").Where(c => c.ReturnValue != 0).

Limitations & notes

  • TTD is user-mode only (a Microsoft limitation): kernel debugging and TTD are distinct session types — you can't time-travel a kernel target.
  • launch and attach_process stop the target at its initial/break-in point with a live process/thread context. (The binding enables the initial-breakpoint event filter — sxe ibp — which a bare DebugCreate host leaves off; without it the target would run free.) Note: on Windows 11 notepad is a Store app, so attaching by the PID that Start-Process notepad returns can hit 0xD000010A (that PID is a transient launcher) — attach to a classic Win32 process.
  • read_memory takes a numeric/0x-hex address only; for register/symbol expressions use execute with db/dd (e.g. db @rip).
  • reachable_from_dispatch is a static call-graph walk over uf disassembly: it follows direct calls and cross-function tail jumps but not indirect calls through function pointers or unresolved compiler jump tables. So a REACHABLE verdict is sound (a concrete static path exists, and the path is reported), while NOT REACHABLE is best-effort within the explored bounds. If the dispatch uses a switch(IoControlCode) jump table (common), pass the specific handler VA as from to scope past it, or confirm dynamically with a breakpoint + go.
  • Crash-dump triage uses !ext.analyze -v, not !analyze — the bundled engine only resolves the module-qualified form (see Bundling the WinDbg engine). On a partial minidump, reads of pages that weren't captured raise An unexpected exception was raised (0x80040205) rather than a clean "memory read error"; query the specific field you need (e.g. dt nt!_DRIVER_OBJECT <addr> DriverName) instead of dumping whole structures. See the crash-dump walkthrough.
  • Single-stepping is only valid once the target is stopped with a real thread context (after a go/step or a breakpoint hit). Stepping straight after a bare goto_position to the very start of a trace (before any thread is live) returns 0x80040205go to a breakpoint first.
  • Symbol names (module!func) need (a) msdia140.dll bundled next to the binary, (b) a symbol path pointing at the PDBs (.sympath …), and (c) for TTD, reloading at a stopped position (after a go/breakpoint, not straight off a !tt) so the module's PDB is matched and loaded. With those, e.g. ttd_calls("ucrtbase!__stdio_common_vfprintf") returns the exact call count. Without symbols, the data model, navigation, and memory reads still work — query by address.
  • One command at a time per session. Sessions run concurrently, but each one is a single engine processing operations serially, so a call against a busy session waits its turn. Issue tool calls for a given session sequentially — await each result before sending the next: the server does not order concurrent in-flight requests against each other, so pipelining them (firing several calls before their results return) can run a command before the one that establishes its state, and is meaningless for stateful, order-dependent debugger commands anyway. Standard MCP clients already serialize call→result, so this is only a concern for custom/batched callers.
  • Sessions are a bounded resource. Each is a process holding a dump, a trace, or a live target, and at most 4 are open at once. end_session when you are done with one rather than relying on the oldest idle session being reclaimed.
  • TTD replay (open_trace) needs TTDReplay.dll discoverable but not elevation; TTD recording (record_trace) needs TTD.exe and Administrator. record_trace captures the recorder's startup output to <out_dir>\ttd_record.log and watches it briefly, so a fast failure (e.g. running un-elevated → 0x80070005 Access is denied) is reported as an error rather than a false "recording started".
  • Control-flow tools (go/step*) issue the corresponding debugger command; precise stop/wait semantics for long-running go against a live target are bounded by the per-call timeout.

Rendered live from glslang/windbg-mcp's GitHub README — not stored, always reflects the source repo.

1 Install Method

NameDescriptionCategorySource
mcpb packageInstall via mcpb (stdio transport)mcp-serverhttps://github.com/glslang/windbg-mcp/releases/download/v0.4.0/windbg-mcp-v0.4.0-windows-x64.mcpb

0 Comments

Login required
Log in to post a comment or update on this repo.

No comments yet — be the first to share an update.