agentic-engineering
A spec-driven workflow for building with coding agents. Never go prompt → code: shape → decide → execute → measure → eval.
Install · Quickstart · What you get · Plugin reference · Claude Code docs
shape decide execute verify diagnose
┌─────────┐ ┌─────────┐ ┌──────────┐ ┌───────────┐ ┌──────────┐
│ /pitch │──▶│ /adr │──▶ │ /ship │──▶ │ /measure │──▶ │ /eval │
│ the spec│ │ the why │ │ the cycle│ │ the data │ │ the modes│
└─────────┘ └─────────┘ └──────────┘ └───────────┘ └────┬─────┘
▲ │ │
│ ▼ │
│ ┌─────────┐ │
│ │ /graph │ fan breadth work out across a fleet, │
│ │ fan-out │ verify each finding, converge │
│ └─────────┘ │
└─────────────── flip-criteria · real failures ◀────────────┘
readers fan out · one writer at a time · review is a gate · data decides
Contents
- Why this exists
- Install
- Staying updated
- Quickstart
- What you get
- Requirements
- Customize per repo
- Maintaining
- Also in this marketplace
- Contributing
- License
Why this exists
Coding agents make it cheap to go from a prompt straight to a diff — and that's exactly the trap. Unshaped work balloons in scope, decisions get lost, and a green build gets mistaken for a correct one.
This plugin encodes a different discipline, distilled from a real, opinionated agentic-engineering practice: shape the work into a written spec first, record the hard decisions, delegate execution to a single serialized writer with an adversarial reviewer as a gate, fan breadth work out across a fleet instead of chaining it one task at a time, and let real data — not a passing test — decide what ships. Six skills, five subagent roles, and verification hooks make that discipline portable to any repo — see the graph model for the underlying lens.
[!NOTE] These are early, opinionated tools — structurally validated and dogfooded on this repo, not yet battle-tested across many. Feedback and issues are welcome. Distributed as a Claude Code marketplace named
giusto-agentic; one marketplace can host many plugins.
Install
In Claude Code:
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace
/plugin install agentic-engineering@giusto-agentic
CLI equivalent
claude plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace
claude plugin install agentic-engineering@giusto-agentic
The six commands and five agents become available. New components load on your next Claude Code session.
Staying updated
Third-party marketplaces don't auto-update by default (a deliberate safety choice — installed plugins shouldn't change under you silently). Two ways to stay current with new releases:
- Hands-off (recommended). Enable auto-update for this marketplace once:
/plugin→ Marketplaces →giusto-agentic→ toggle auto-update on. Claude Code then refreshes it on startup and, when a new version ships, notifies you to run/reload-plugins— no manual checking. - Manual. Pull the latest yourself:
then/plugin marketplace update giusto-agentic /plugin update agentic-engineering@giusto-agentic # and/or github-keeper@…/reload-plugins(or restart). The CLI equivalents areclaude plugin ….
Releases are versioned — see the releases and the CHANGELOG.
Quickstart
A typical end-to-end cycle, from idea to merged code:
- Shape it.
/pitch "add Google OAuth to login"— an interview producesdocs/pitches/oauth.md, naming the graph the work touches and its gated edges. Approve the shape before any code is written. - Decide it.
/clear, then/adrif a consequential choice was made (e.g. session strategy). Behavior-altering work records aFlip-criteria. - Fan it out (if it has breadth).
/clear, then/graphfor audits, reviews, or research that spans many independent pieces — fan out across a fleet of subagents, verify each finding, converge. - Ship it.
/clear, then/ship oauth— theexecutorimplements and opens a PR; thereviewerchecks the diff against the pitch in a fresh context; you fix the real gaps and merge. - Prove it. Behavior-altering changes ship dark (flag-OFF), then
/measureagainst the recorded criterion flips them on — on real data, not a green build.
[!TIP] Run
/clearbetween stages. Separating exploration, decision, and execution into clean contexts is the single biggest lever on output quality.
What you get
Skills — slash commands
| Command | What it does |
|---|---|
/pitch | Shape a feature into a Shape Up pitch (the spec / source of truth) via interview, before any code. Names the graph the work touches and its gated edges. |
/adr | Record a consequential decision in an append-only docs/DECISIONS.md, with optional dark-launch Gate/Flip-criteria. |
/graph | Turn a straight-line task into an execution graph: fan breadth work (audits, reviews, research) out across a fleet of subagents, verify findings, converge. Built on dynamic workflows — coordination costs zero model tokens. |
/ship | Execute an approved pitch as a closed-scope cycle: right-size gate (is the ceremony warranted at all?), pre-spawn filter, doc-bundle, standard PR format, adversarial review, dark-launch flip. You invoke it — a cycle spawns a writer and opens a PR, so the model never starts one on its own. |
/measure | Unblock a decision with a read-only, data-backed flip/keep/cut verdict — never guesses, never writes. |
/eval | Make eval the unit of progress: build the harness from real failures, localize where a pipeline breaks (transition-failure matrix), feed flip-criteria. |
Agents — roles
The organizing law is parallelize readers, serialize writers: research and review fan out across many agents; only one writer ever touches the code.
| Agent | Role |
|---|---|
executor | The serialized writer: implements one closed scope, keeps build/tests green, opens a PR, never merges. |
researcher | Fan-out, read-only research with cited, verified findings; prefers current docs over memory. |
reviewer | Adversarial pre-merge gate over a whole PR diff; 8-check rubric; returns MERGE / ADJUST / REJECT. |
verifier | General single-claim adversarial node: given one finding, tries to kill it and returns real/not-real. Fan-out safe — run several in parallel as a gate on a graph edge. |
measurer | Read-only data verdicts to unblock measure-gated decisions. |
Hooks
- PostToolUse — auto-format/lint edited TS/Py files (only if
eslint/ruffare present). - PostToolUse — verify
path:lineclaims written into markdown: afile:lineis the kind of claim a reviewer nods at and a command settles, so a command settles it. It reports only references whose path resolves on disk and whose line is out of range or blank — a clean run is silent, a report is never a guess, and paths that don't resolve (examples, other repos, planned files) are skipped. Needspython3; silently inert without it. - Stop — an advisory reminder: when a turn leaves uncommitted changes, it prints a one-line nudge to run the project's own checks before declaring done. It never blocks and never re-invokes the model (so it can't loop, even with background workflows); it's silent when the tree is clean.
[!NOTE] The Stop hook is a reminder, not a gate — it surfaces "you have uncommitted changes; run this project's checks and show the output" and then gets out of the way. The real verification discipline lives in the
/shipandexecutorprompts; the hook just keeps it front-of-mind without ever blocking the turn.
Requirements
- Claude Code (CLI, desktop, or IDE extension).
jqon your machine, for the format hook.- (optional) A docs MCP such as Context7
so
researcher//measurecan pull live library docs:claude mcp add context7 -- npx -y @upstash/context7-mcp.
Customize per repo
The skills reference docs/pitches/ and docs/DECISIONS.md by convention. Adapt
the doc-bundle tiers in
EXECUTION_PLAYBOOK.md
and the invariants in the review checklist to your project, and change the paths
if your repo differs. See the plugin reference
for the full component breakdown.
Maintaining
Releasing an update: bump version in both
plugin.json and
marketplace.json, commit, tag, and push.
Users pick it up with /plugin marketplace update giusto-agentic.
Validate before pushing:
claude plugin validate .
claude plugin validate ./plugins/agentic-engineering
Repo layout
agentic-engineering-marketplace/
├── .claude-plugin/
│ └── marketplace.json # lists the plugin(s), schema-validated
├── assets/
│ └── banner.svg
├── docs/
│ ├── DECISIONS.md # append-only ADR log
│ └── pitches/ # shaped specs, one per feature
├── plugins/
│ ├── agentic-engineering/ # the spec-driven / graph-engineering plugin
│ │ ├── .claude-plugin/plugin.json
│ │ └── skills/ agents/ hooks/ README.md
│ └── github-keeper/ # README + open-source maintenance plugin
│ ├── .claude-plugin/plugin.json
│ └── skills/ README.md
├── LICENSE
└── README.md
Also in this marketplace
github-keeper — make a public repo
well-made and honest. /readme's fidelity gate derives the hero archetype,
palette, voice, badges, and sections from the target project's own identity and
proposes-and-confirms before generating, instead of imposing a house style;
/opensource adds the community-health files, an honest CI gate, and the right
repo settings. Install it with /plugin install github-keeper@giusto-agentic.
Contributing
Issues and PRs are welcome — bug reports, a sharper skill prompt, or a new
role/skill that fits the shape → decide → execute → measure → eval spine.
Open an issue
to discuss anything non-trivial first. Run claude plugin validate . before
opening a PR.
License
MIT © Giusto Piedimonte