Back to Discover

Writ

plugin

infinri

Writ keeps AI coding on your rules. It feeds Claude Code the right coding standards for each task in under a millisecond, and it blocks risky code changes until a human approves the plan and the tests. Ships with 287 rules across security, architecture, and testing.

View on GitHub
164 starsMITSynced Aug 2, 2026

Install to Claude Code

/plugin marketplace add infinri/Writ

README

Writ

A Claude Code harness that gives every coding session two helpers: a fast librarian that picks the rules that fit the current task, and a process keeper that blocks risky writes until you have approved a plan and tests.

The librarian returns ranked results in 0.6 ms at the 95th percentile (measured 2026-08-01 at the live 287-rule corpus). At the 10,000-rule synthetic scale it still holds at 0.83 ms while reducing context tokens by 749 times versus loading the whole rulebook every turn.

See CHANGELOG.md for the release history through v1.6.0 (re-measured benchmarks with environment disclosure, hook-system audit and hardening, force-swap coverage, the Claude Code 2.1.220 black-box refresh).

Browse the architecture in your browser

Six self-contained HTML pages render the whole system with diagrams, served live on GitHub Pages: system overview, data model, retrieval pipeline, injection channels, knowledge graph explorer, and corpus round-trip. (The source lives under docs/architecture/; cloning and opening index.html locally works too, no server needed.)

Install as a Claude Code plugin

Writ is published as a single-plugin marketplace in this repo.

Prerequisites: Python 3.11+, Docker (Neo4j runs in a container), jq, curl, envsubst.

claude plugin marketplace add infinri/Writ
claude plugin install writ@writ

Find the install directory (the next command needs it):

WRIT_DIR=$(claude plugin list --json \
  | python3 -c "import json,sys; print(next(p['installPath'] for p in json.load(sys.stdin) if p['id'].split('@')[0] == 'writ'))")

One-time bootstrap. Creates the venv at ${CLAUDE_PLUGIN_DATA:-$HOME/.cache/writ}/.venv, brings up Neo4j, seeds the rule corpus from writ-corpus.cypher, and starts the FastAPI daemon:

bash "$WRIT_DIR/scripts/bootstrap-plugin.sh"

Restart Claude Code and verify with curl http://localhost:8765/health.

Patch global config (plugin mode only). Plugin installs do not write to ~/.claude/settings.json or ~/.claude/CLAUDE.md, so the Bash allowlist that suppresses permission prompts and the workflow instructions are missing until you run this once (idempotent, backs up existing files, --dry-run to preview):

bash "$WRIT_DIR/scripts/patch-global-config.sh"

The plugin's hooks degrade gracefully until bootstrap completes: sessions are never blocked, and the SessionStart hook prints setup instructions while anything is missing. Full install detail, the standalone path, the systemd service, and troubleshooting live in docs/install.md.

The problem

Three things break when you give a coding agent a large rulebook the obvious way (paste it all into the prompt):

  1. Token cost grows with the rulebook, not the work. At 80 rules: about 13,876 tokens of rule text every turn. At 10,000 rules: 1,174,142 tokens. Cache hit rates collapse, latency climbs, the bill scales with the rulebook.
  2. Relevance degrades. A model handed every rule treats none of them as load bearing. Specific rules drown in generic ones.
  3. Workflow discipline has nowhere to live. Static skill files can describe a process; they cannot enforce it. Telling the model to write tests first does not stop it from writing the implementation first.

What Writ does about it

Two layers, sharing a Neo4j-backed knowledge graph:

The knowledge layer (the librarian). A FastAPI service on localhost:8765 that runs a five-stage hybrid retrieval pipeline:

Query text
  Stage 1: Candidate filter          < 1 ms    category-routed or domain-scoped corpus
  Stage 2: BM25 keyword (Tantivy)    < 2 ms    top 50 candidates
  Stage 3: ANN vector (hnswlib)      < 3 ms    top 10 candidates
  Stage 4: Graph traversal           < 3 ms    adjacency cache for DEPENDS_ON,
                                                SUPPLEMENTS, CONFLICTS_WITH, ...
  Stage 5: Two-pass ranking          < 1 ms    reciprocal rank + context budget
                                              total p95 budget: 10 ms

Each retriever covers a blind spot the others have. BM25 catches exact keyword matches. Vectors catch paraphrase ("SQL" versus "database query"). Graph traversal catches rules that share neither but are causally related. The two-pass ranker fuses everything with severity, confidence, and graph-proximity weights, and an abstention gate returns nothing rather than injecting noise when the best raw cosine falls below 0.30.

The enforcement layer (the process keeper). 37 hook scripts under hooks/scripts/, wired into Claude Code via hooks/hooks.json (41 registrations across 12 hook events), plus one statusLine script, a session state machine in writ/session/, slash commands, and 5 sub-agent role files under agents/. The state machine owns mode, phase, and gate state; hooks are thin clients that delegate to it.

Mandatory rules (the architectural invariant). Rules with mandatory: true (33 in the live corpus, spanning ENF-* enforcement rules and SEC-/PERF-/SCALE-* invariants) are excluded from the retrieval pipeline at index build time. They reach the agent out of band through the /always-on endpoint with its own 5,000-token budget, enforced by a corpus integrity check. No change to ranking weights, embedding model, BM25 tuning, or graph traversal can cause an enforcement rule to disappear from agent context.

Quick start (standalone)

git clone <writ-repo> ~/.claude/skills/writ
cd ~/.claude/skills/writ
bash scripts/bootstrap.sh

The bootstrap script handles everything: prerequisite checks, virtualenv, dependency install, global config rendered into ~/.claude/, rule and agent symlinks, Neo4j via Docker Compose, corpus ingestion, and daemon startup. Idempotent.

Verify:

writ status
# {"status":"healthy","rule_count":287,"mandatory_count":33,"index_state":"warm",...}

writ query "controller contains SQL query"
# Mode: full | Candidates: 14 | Latency: 0.3ms
# 1. [0.984] SEC-INJ-SQL-001  Parameterized queries only...

Open Claude Code in any project. Type a prompt. You should see a [Writ: ...] status line and a --- WRIT RULES --- block with the rules that apply to what you are doing.

Editing rules directly. The corpus ships as writ-corpus.cypher, a single portable dump, not as a tracked bible/ directory. To hand-edit a rule instead of going through writ add/writ edit: writ export materializes bible/ locally from the live graph, edit the markdown, writ import-markdown pushes it back into Neo4j, then writ export-cypher writ-corpus.cypher refreshes the shipped dump before committing.

What you experience

When Claude is doing read-only work (asking questions, debugging, reviewing), Writ injects relevant rules and stays out of the way. When Claude is in Work mode and tries to write code before you have approved a plan, the write is denied with a clear reason like:

[ENF-GATE-PLAN] Write blocked. Approve plan.md first.

You write the plan, you say "approved," the gate opens. The next gate (test skeletons) blocks code that has no tests pointing at it. Same pattern: write the tests, approve, gate opens. After both gates clear, Claude writes the implementation freely.

Approval cannot be self-served. Advancing a gate requires a one-time token written to /tmp/writ-gate-token-${SESSION_ID} only when the user actually types an approval phrase. If Claude tries to advance the gate itself, the token is missing and the call is denied (logged as agent_self_approval_blocked).

The mode system

ModePurposeGating
ConversationDiscussion, brainstorming, questionsNone
InvestigateEvidence-grounded audit, exploration, researchCoverage and citation tracking; web research requires two independent sources before synthesis
DebugInvestigating a specific problemSource edits blocked until debug.md records a root cause
ReviewEvaluating code against rulesNone
WorkBuilding or modifying codeTwo approval gates before implementation

In Work mode, two gates apply:

  1. phase-a validates plan.md against four required sections (## Files, ## Analysis, ## Rules Applied, ## Capabilities) and verifies every cited rule ID against rules actually loaded in the session.
  2. test-skeletons requires at least one test file with real assertions before production code is written.

Decision memory: the repo remembers why files changed

Every commit made under Writ is captured into the same Neo4j graph as the rulebook: the decision behind it (title, rationale, the approved plan's per-file reasons), each file change with the rules the AI was shown and the rules it cited, and the commit itself, linked by typed edges. Capture is a post-commit git hook that never blocks a commit and fails open when the daemon is down; writ harvest backfills history.

That record then plays back in three places:

  • writ recall compiles recent decisions into a token-budgeted digest, and a once-per-session briefing injects the top of it into your first prompt, so a new session starts knowing what was decided and why.
  • writ pr sync posts one comment per changed file on the open Bitbucket pull request: the captured reason for the change, the rules the AI was shown, and the rules it cited. Idempotent (it updates its own comments), so reviewers see the "why" next to the diff instead of reconstructing it.
  • Git notes: the same per-file reasons are written to refs/notes/writ-decisions, a plain-git channel that travels with the repository and needs no server.

This feature family is adapted from concepts pioneered by JolliAI: capturing AI development decisions as durable records, replaying them into new sessions, and pushing them onto commits and open PRs. Writ's recall eviction policy is adapted from Jolli's ContextCompiler (the policy, not the code; see writ/session/recall.py), and adds Writ's own twist: every decision stays grounded in the governing rule IDs, which are never evicted from the digest. Full detail: docs/reference/decision-memory.md.

Performance

Live system measurement (2026-08-01, 287-rule corpus, ONNX runtime, warm indexes; steady-state samples over 10 representative queries):

StageMedianp95BudgetHeadroom at p95
End to end0.4 ms0.6 ms10.0 ms17x

Synthetic scale curve (2026-08-01, from SCALE_BENCHMARK_RESULTS.md):

CorpusE2E p95Tokens stuffedTokens retrievedReduction
80 rules0.383 ms14,7382,0567.2x
500 rules0.486 ms79,4911,89841.9x
1,000 rules0.662 ms137,9901,68681.8x
10,000 rules0.827 ms1,190,6491,590748.8x

All of the above was measured on a 16-thread AMD Ryzen 9 7940HS laptop with 31 GiB RAM and an uncapped Neo4j container (512M pagecache); absolute numbers will differ on smaller machines. The full disclosure is the "Measurement environment" section of SCALE_BENCHMARK_RESULTS.md.

Retrieval quality against the 193-query ground-truth corpus (47 ambiguous, expanded 2026-07-17). The floors are regression gates the build fails below, deliberately set under the measured values, not quality targets:

MetricFloorMeasured (2026-08-01)
MRR at 5 (ambiguous queries, n=47)>= 0.450.5681
Hit rate at 5 (all 193 queries)>= 0.750.7824
Domain hit rate top-5>= 0.900.9323
nDCG at 10>= 0.650.7071
Methodology MRR at 5 (n=40, signed off)>= 0.780.8271
Methodology hit rate>= 0.900.9500

Full numbers in SCALE_BENCHMARK_RESULTS.md. Architectural detail in HANDBOOK.md.

Relationship to Agent Skills

Anthropic's Agent Skills standard, originated by Anthropic and now maintained as an open spec at agentskills.io, solves a specific problem well. Skills live as directories with a SKILL.md entry point; at session start the agent pre-loads each skill's name and description into its system prompt. This is progressive disclosure: the agent sees just enough to know when a skill applies, without the body of the skill consuming context until needed. The design optimizes for a small system-prompt footprint, agent-side relevance decisions, no per-skill load cost, and low-friction authoring.

Writ targets a different problem. Writ is built around an enforcement-grade rule corpus, currently 287 rules and designed to scale into the thousands, where:

  • The agent must not be the matching-decision-maker between context and rule.
  • Retrieval must be triggerable by filesystem and tool-call signals, not only by prompt content.
  • The corpus exceeds what fits in pre-loaded descriptions even with progressive disclosure (1,190,649 tokens at 10,000 rules versus a pre-loaded budget measured in low thousands).

Where progressive disclosure runs out

The pattern works as long as the agent's matching of descriptions to context is reliable. The matching step degrades as:

  • Skill count grows. Each pre-loaded description costs system-prompt tokens; with hundreds of entries, descriptions blur and discrimination drops.
  • Descriptions overlap. Two skills with semantically nearby triggers compete; the agent picks one and silently skips the other.
  • Context becomes ambiguous. A prompt that touches several domains gives the agent multiple plausible matches; pre-loaded descriptions don't disambiguate.
  • The trigger isn't in the prompt. A skill that needs to fire when the user opens a Python file, runs a specific command, or modifies a config can't be matched from prompt text alone.

These are not bugs in the spec. They are the boundary of what agent-side matching against pre-loaded text can do.

Writ's alternative

Writ models skills, playbooks, rationalizations, and forbidden-response sets as nodes in a hybrid-RAG knowledge graph (Neo4j-backed), retrieved by the same five-stage pipeline that surfaces rules: candidate filter, BM25 keyword (Tantivy), ANN vector (hnswlib), graph traversal over a pre-computed adjacency cache, weighted ranking. The agent does not match descriptions; the pipeline matches text plus tool state plus filesystem context.

Two architectural splits make this practical:

  • Retrieval-on-demand for the bulk of the corpus. Ranked results in 0.6 ms at p95 at the live corpus; at 10,000 rules, 0.83 ms p95. Retrieved tokens stay roughly flat (around 1,600-2,000) regardless of corpus size, while context stuffing scales linearly (14,738 tokens at 80 rules to 1,190,649 at 10,000). Reduction: 56x at the live corpus, 749x at 10,000 rules (measured 2026-08-01).
  • Always-on bundle for the mandatory floor. Mandatory rules and forbidden-response nodes load every turn through a dedicated endpoint with its own 5,000-token budget. They are excluded from the retrieval pipeline at index build time, so no ranking change can cause an enforcement rule to drop out of agent context.

Thirty-seven hook scripts read filesystem changes, tool calls, and session state, and invoke retrieval with those signals attached. Methodology arrives in the agent's context based on observable state, not on whether the agent recognized the trigger from prompt text alone.

Boundary

Agent Skills is the right answer for small skill counts, human-authored discrete behaviors, contexts where agent-side matching against descriptions is acceptable, and authoring workflows that prioritize zero-config installation.

Writ's model is the right answer for large enforcement-grade rule corpora, contexts where the matching decision must move out of the agent, and pipelines that must fire on tool calls and filesystem state, not just prompts.

Same problem space, different optimization frontiers.

CLI reference

The most-used commands; run writ --help for the complete list.

CommandWhat it does
writ serve [--host --port]Start the FastAPI service (default localhost:8765).
writ statusHealth check via the HTTP service.
writ query <text> [--domain --budget --local]Run a retrieval query (server first, in-process fallback).
writ import-cypher [input]Rebuild the graph from the Cypher dump (wipes first). Default writ-corpus.cypher; this is what install and CI use.
writ export-cypher [output]Dump the whole graph as the portable, tracked Cypher script.
writ import-markdown [path] [--only --dry-run --compress]Ingest bible/ markdown (rules + methodology) into Neo4j.
writ export [output]Regenerate bible/ markdown from the graph.
writ reconcile [--bible-dir --project]Delete graph nodes/edges/props absent from source (the only prune path).
writ validate [--benchmark]Run the ~30 corpus integrity checks.
writ add / writ edit <rule_id>Interactive rule authoring with schema, redundancy, and conflict checks.
writ propose ... / writ review [rule_id]AI rule proposal through the structural gate; human triage.
writ compressCluster rules into Abstraction nodes for summary mode.
writ doctor [--fix --net]13 install health checks; --fix repairs 6 of them.
writ logs [tail|stats|list|rotate|backup]Read and manage the typed log streams.
writ harvest [--since] / writ recallDecision-memory backfill from git + transcripts; recall briefing.
writ pr syncPost captured file-change reasons as Bitbucket PR comments.
writ analyze-friction [flags] / writ audit-session <sid>Friction-log analytics and per-session timelines.

API reference

All endpoints under http://localhost:8765. JSON bodies; no auth (binds localhost only). Total: 45 endpoints: 11 query/corpus routes, 3 gate routes, 24 session-state routes under /session/{id}/, 2 decision-memory, 1 git-hooks, 4 explorer.

MethodPathPurpose
POST/queryRun the 5-stage pipeline. Returns {rules, mode, total_candidates, latency_ms, abstain_signal}.
POST/prompt-bundleOne warm call combining broad query + always-on + methodology companion (the per-prompt hot path).
GET/always-onMandatory rules plus ForbiddenResponse nodes plus always-on Skills/Playbooks; mode- and injection-point-scoped.
POST/methodology-companionDeterministic workflow-state methodology matching (floor/push/pull channels).
POST/analyzePattern analysis of a code snippet (LLM escalation only if the anthropic SDK is installed).
GET/rule/{rule_id}Fetch a rule, optionally with one-hop graph context.
POST/propose / POST /feedback / POST /conflictsRule proposal gate, feedback signals, conflict checks.
GET/healthNeo4j round-trip; rule count, mandatory count, index state, cache dir.
POST/pre-write-checkConsolidated write gate + escalation + file-context RAG.
POST/session/{sid}/advance-phaseToken-gated gate advance (the anti-self-approval surface).
GET/dashboard / GET /explore / GET /graphFriction analytics HTML and the graph explorer.

Errors come back as HTTP 200 with {"error": "..."} for logical failures, 422 for Pydantic validation, 5xx for unhandled exceptions. Clients should check the error key, not just status codes.

Configuration

FilePurpose
writ.toml[neo4j] credentials, [hnsw] cache dir, [bitbucket] credentials, [logs] backup destination. Gitignored; writ.toml.example is the template. Missing file falls back to dev defaults.
pyproject.tomlPackage metadata. Production deps: fastapi, uvicorn, neo4j, tantivy, hnswlib, httpx, pydantic, typer, rich, onnxruntime. sentence-transformers is the optional [fallback] extra, not a production dep.
.claude-plugin/plugin.jsonPlugin manifest. Deliberately declares no hooks key (auto-discovery of hooks/hooks.json; declaring it collides) and no agents key (declaring one loads zero agents; auto-discovery of agents/ loads all 5).
.claude-plugin/marketplace.jsonSingle-plugin marketplace so claude plugin marketplace add infinri/Writ resolves. Version must match plugin.json.
docker-compose.ymlSingle neo4j:5 service on 7474/7687, health-checked. The daemon itself runs natively.
hooks/hooks.jsonCanonical hook wiring: 41 registrations across 12 events over 37 scripts. Auto-discovered; single source for the generated templates/settings.json.
bin/lib/gate-categories.jsonGate exclusion globs plus framework detection.
writ/shared/budget.jsonBudget constants: default 8000, rule costs 200/120/40, always-on cap 5000.

Common environment variables: WRIT_HOST/WRIT_PORT (daemon target), WRIT_CACHE_DIR (session caches, default <install>/var/session), WRIT_LOG_ROOT (typed log streams, default <install>/var/logs), WRIT_FRICTION_LOG (collapse all streams to one file), WRIT_DEBUG (debug sinks, default off), WRIT_NO_AUTOSTART (suppress daemon autostart), WRIT_ALLOW_EMBEDDING_FALLBACK=1 (permit the sentence-transformers path when ONNX is absent). Neo4j credentials are read from writ.toml only.

Testing

367 test modules, ~5,700 collected tests. Roughly half exercise a live Neo4j and skip only when the graph is genuinely unreachable; an empty-but-reachable graph fails loudly instead of skipping. The suite runs against a dedicated daemon port (8799) and isolated cache/log directories, and restores the production corpus from writ-corpus.cypher when it finishes.

make test          # pytest tests/ -x -q
make bench         # pytest benchmarks/bench_targets.py -x -q
make check         # both, plus writ validate

Pre-commit: make bench runs at pre-push.

Evidence: pressure runs and operational reviews

For auditors (and anyone deciding whether to trust an enforcement tool), Writ keeps its verification artifacts in the repository rather than asserting them:

  • docs/pressure-runs/: adversarial scenario runs against real Claude Code sessions. Each completed run ships the verbatim task prompt, the full session transcript, the friction-log delta (every hook decision as JSONL), and a graded analysis.md scoring each targeted rule as held / bypassed / unclear. PSR-008 is a complete example: 9 scored criteria, 8 passed, one designed-in failure documented honestly, and two real findings filed from the run. The methodology is in the runbook.
  • docs/monthly-reviews/: recurring operational reviews built from the friction log via writ analyze-friction: per-rule activation counts, denial stick rates, rationalization counts, skill usage, and playbook compliance, with keep/revise/trim actions per rule. 2026-05.md is the first completed review (7,058 logged events in the window).

Both artifact sets are produced by the system's own audit stream, not written after the fact.

Status

v1.6.0 (2026-08-01). Every benchmark re-measured on disclosed hardware; documentation rebuilt from a full code read; the 37-script hook system audited end to end (silent-failure fixes, force-swap coverage, fail-open posture documented as the specification); Claude Code contract re-pinned to 2.1.220. Installs end to end as a Claude Code plugin (verified Agents (5) on 2.1.220). Every number in this README is either measured and dated, or derived from the current source tree.

Related documents

  • HANDBOOK.md: the operator manual (modes, gates, approval model, sub-agents, corpus).
  • docs/install.md: full install detail for both install paths, the systemd service, and troubleshooting.
  • SCALE_BENCHMARK_RESULTS.md: the full live measurement plus the synthetic scale curve.
  • CONTRIBUTING.md: rule authoring workflow, review cadence, AI proposal triage.

Acknowledgements

Jesse Vincent's Superpowers revealed gaps in Writ's methodology coverage, and observing that project's design choices in practice fed Writ's analysis of the Agent Skills format (see Relationship to Agent Skills).

The decision-memory feature family (capturing decisions to the graph, session recall, and pushing per-file reasons onto commits and open PRs) is adapted from concepts pioneered by JolliAI; the recall eviction policy in particular is adapted from Jolli's ContextCompiler (see Decision memory).

License: MIT. Authored by Lucio Saldivar.

Rendered live from infinri/Writ's GitHub README — not stored, always reflects the source repo.

1 Plugin

NameDescriptionCategorySource
writHybrid RAG rule retrieval + session-aware workflow enforcement. Injects relevant coding rules via a 5-stage pipeline (BM25 + vector + graph traversal + reciprocal-rank fusion) and enforces mode-based workflow gates (plan -> tests -> implementation).workflow./

0 Comments

Login required
Log in to post a comment or update on this repo.

No comments yet — be the first to share an update.