Claude Code toolkit · Render · Open source · formerly Shipwright
Consul: Turn a Plain-Language Spec Into a Live App
A Claude Code toolkit that takes a plain-language description of an app and sees it through to live software. Clone the repo, type /start, and answer four questions. A new monorepo appears on GitHub with its database and two web services already running on Render. Then /spec interviews you and writes a brief, one card with requirements and acceptance criteria, and a single /goal line runs it. A runner delegates to three worker agents, a checker model judges when the brief is finished, and hooks make the bookkeeping mandatory.
TL;DR
- What it is. Nine skills, three worker agents, three hooks, and a set of scripts that install into a new project. /start creates the project, /spec writes the brief, and /goal runs it to the end.
- The contract. Clone, type /start, and a new repo exists with its services live on Render. Nothing per-project may sit between those steps, so a provisioner script creates the cloud resources and the only manual setup is one-time account work.
- The brief. A markdown card on a committed briefs/ board where the folder is the status, 1-backlog, 2-active, 3-blocked, 4-done. It carries its own requirements, execution protocol, task breakdown, progress log, and blockers.
- Two endings. Human blockers are parked while everything else keeps running. A brief lands in 3-blocked only when nothing runnable remains, lands in 4-done only when every requirement is verified, and a blocked brief is never reported as done.
- Enforced by the harness. /goal keeps turns coming and an independent checker decides completion. A Stop hook refuses to end a turn when the brief was left untouched, and another hook logs every worker call.
- The proof. In the first end-to-end build, a four-requirement brief went from backlog to 4-done in one runner turn. About 42 minutes, 7 subtasks, 0 blockers, and 11 backend tests passing against a real Render Postgres database.
01Three steps and a contract
Claude Code can write a working app in an afternoon. Everything around the code is where a person who does not write software stalls. Someone has to create a repository, a database, and a host, wire credentials between them, track what is finished across sessions, and decide when the thing is done. Consul is that surrounding layer. I started it in March 2026 as Software Factory, renamed it Shipwright in September, and eight days later renamed it Consul. The machinery stayed the same through all three names.
The repository states its promise as a contract with three steps. Clone the repo. Open Claude Code in it and type /start. A new project repo exists, one monorepo with backend/ and frontend/, pushed to GitHub, with its Render database and web services live and Claude Code configured inside it. Claude Code is the only thing a person must have installed. /start installs whatever tools are missing, walks through GitHub and Render one click at a time, asks four questions, and runs the setup wizard.
Every change to Consul gets measured against those three steps. Nothing may become a per-project manual step between them. No Dashboard clicking, no file editing, no "first go create X." The only manual setup allowed is account-level and done once. Add a payment method on Render, let Render read your GitHub, create an API key, and optionally create a Clerk app if people will sign in. If a feature needs something from the human, the wizard asks for it or the readiness check reports it with a fix line. If a script can do it, a script does it. A small project runs about $20 a month on Render while it exists, and one command deletes it.
02A brief on a board
Inside the new project, /spec create interviews you about the product. It reads the project's CLAUDE.md first, so it already knows the stack and deploy target, and asks only product questions. What are you building, who is it for, what must it do, what should it leave out, and how will you know it is done. Before drafting, it runs one check on itself. Could a fresh session break every requirement into subtasks without asking anything? If the answer is no, it asks the follow-up now, while you are still there.
The output is a brief, one markdown file on a committed board called briefs/. The board has four folders, 1-backlog, 2-active, 3-blocked, and 4-done, and the folder a brief sits in is its status. There is no status field to drift out of sync. The brief carries everything a session needs to pick up the work cold. Each requirement has a permanent ID and acceptance criteria written as something checkable, a command output, a visible screen state, or a status code. An execution protocol binds any session that works the brief. Below that sit a task breakdown the runner fills in, an append-only progress log, and a blockers section where each question waits with its options and an empty Resolution: line.
POST /api/clients returns 201 and the client appears on the Clients screen.PATCH sets status paid and the invoice moves to the Paid column.Type: needs-human-decision · Subtask: p2-task-2
Description: The outstanding total needs one currency or a rule for mixing them.
Options: 1. USD only for now 2. A currency per client, totals grouped by currency
Resolution: USD only for now. (answered in chat, written here by the runner)
Two details carry a lot of weight. The work-in-progress limit on 2-active is one, so a session never has to wonder which brief it is working. And the brief is the only state. Close Claude Code mid-run, reopen it, type /orchestrate, and the runner reads the board and continues where the progress log leaves off.
The spec skill finishes by handing you a /goal line to paste. Its condition is a location. The brief sits in 4-done or 3-blocked, and the prompt names them as two distinct terminals. Every turn has to prove the board state by listing the folders, and the turn cap is three times the requirement count plus ten, so a four-requirement brief gets 22 turns. Long runs are fine. Endless ones are impossible.
03One runner, three workers
Pasting the /goal line starts the runner, a skill called orchestrate running in the main session. Its first paragraph sets the job. It plans, spawns, verifies, records, and routes, and it must not write implementation code. It pulls the next brief into 2-active, runs the readiness check, and groups the requirements into phases by dependency. Infrastructure comes first, then the backend with its deploy, then the frontend with its deploy. Only the current phase is broken into subtasks. Later phases wait as placeholder rows until the earlier workers' interface contracts are in hand.
Each subtask names a worker, the requirement IDs it advances, what it depends on, and an exclusive list of files it owns. No two subtasks in a phase may own the same file. The three workers are native Claude Code subagents with their skills preloaded through frontmatter. The infra-worker carries the deploy skill and applies render.yaml through the provisioner. The backend-worker carries backend-test and runs its tests against the real Render Postgres database, reached through credentials in backend/.env, with no mocks and no SQLite stand-in. The frontend-worker carries bold-design, which pushes the interface toward the product's own world, and verify-ui, which screenshots the page and iterates until it matches the requirement.
Workers run one at a time through the Task tool. Each starts with fresh context and ends with a fixed report, a STATUS of completed, blocked, or failed plus evidence, files, interface contracts, and decisions. The runner spot-checks the evidence before it believes a report, and it checks a requirement's box only when every subtask behind it is complete and the evidence meets the acceptance criteria. A failed subtask is respawned with the specific failure, up to three attempts. The third failure becomes a blocker.
The first version spawned each worker as a headless claude -p process with an --agent flag and a $50 budget cap per task. Native subagents replaced all of that machinery, and the /goal turn cap became the single backstop. Pushing belongs to the runner. Under the human push policy, a finished deploy subtask becomes a blocker asking you to run git push. Under the consul policy, the runner pushes, logs the commit SHA, and spawns a verification subtask. Workers never push.
04Blocked is its own ending
The rule I care most about in Consul is the line between blocked and done. An autonomous run that reaches a question it cannot answer has two bad options. It can stop and wait, leaving hours of unrelated work undone. Or it can guess, and report success on something that needed a decision. Consul does neither.
When a worker comes back blocked and the runner cannot clear it alone, it records the blocker in the brief with a type, a description, the context, the options, and an empty Resolution: line. Then it recomputes the runnable set and keeps going. A subtask is runnable when it is pending, its dependencies are complete, and it sits downstream of no blocked subtask, whether by dependency, by a shared file, or by a shared requirement. The blocker waits in the brief while everything else moves.
Routing follows requirement state. When every requirement box is checked, the brief moves to 4-done with outcome completed. When requirements remain and zero runnable subtasks remain, it moves to 3-blocked with outcome needs-human, and the runner announces NEEDS HUMAN INTERVENTION with every open question and its options. Anything else stays in 2-active for another turn. A logged blocker that ended up gating no requirement rides into 4-done as a follow-up note. The two terminals never blur, and the /goal prompt itself forbids a blocked brief from claiming success.
The router below runs those rules. Set each requirement's state, say whether other runnable work remains, and watch where the brief goes.
Answering is short. You fill in the Resolution: line, or in Quick Start you answer in chat and the runner writes it into the brief for you. The next /orchestrate sees the filled line, moves the brief back to 2-active, marks each answered blocker resolved, and re-plans only the subtasks that were waiting on it.
05What the harness guarantees
A long autonomous run cannot depend on the model remembering its instructions on turn fifteen. Consul moves the rules that matter out of the skill text and into the harness.
- The loop. /goal keeps turns coming until its condition is met, and an independent checker model decides whether it is. The runner ends every turn by listing the board folders, so the checker judges from the listing and the runner's own summary counts for nothing.
- The progress log. A UserPromptSubmit hook touches a marker file when each turn starts. A Stop hook compares the active brief's modification time against that marker, and if the brief was left untouched it exits with code 2. The turn cannot end until a progress log entry is written.
- The trajectory. A PostToolUse hook on every Task call appends one JSON line to
session/{brief-id}/trajectory.jsonlwith the agent type, the description, and the prompt and response sizes. It always exits 0, because logging must never break a run. The model writes a curatedtrajectory.mdbeside it, and the JSONL is the record evals can rely on. - The workspace. A PreToolUse hook on Bash inspects every Render CLI command and API call and blocks it unless the current workspace matches the project's pin, by name or by ID. It fails closed. When the CLI cannot answer, it asks the API which owners the key can see.
Single writer is the last rule and the simplest. Only the runner edits the brief or moves it between folders. Workers read it and never write it. A board with one writer never has to reconcile two versions of the truth.
06Scripts where people used to click
For most of the project's life, one manual step sat between the wizard and the first run. The human created a Render Blueprint Instance in the Dashboard, because the runner was forbidden from creating resources through the API in case it hit the wrong workspace. A non-developer stalls exactly there. Once the workspace pin and an explicit owner ID made the wrong-workspace failure impossible, the rule had no reason left, and a script took over the step.
provision.py applies render.yaml through the Render API. It creates env groups, creates databases and waits until they are available, creates services from the GitHub remote, fills in cross-service URLs and secrets, and writes backend/.env plus a file of live service IDs and URLs. It is idempotent by name, pinned by owner ID, and never deletes on its own. It runs at the end of onboarding and again as the infra-worker's first task. Its counterpart, provision.py --destroy, lists what it will remove, asks you to type the project's name, and removes it. A smoke test that leaves paid resources behind is a failed smoke test.
The readiness check follows the same idea. preflight.py is standard-library Python. It checks tools, the git remote, the Render credential and its expiry, the workspace pin, and render.yaml against the live services, database, env group, secrets, and health. Each FAIL line comes with a fix written for someone who has never opened a terminal. The runner runs it before starting a brief, and every Render FAIL becomes a blocker carrying the fix text verbatim. In the first live smoke test of the contract, the database took about two and a half minutes to become available, both services went live within about a minute of being created, and 17 preflight checks passed on the live path.
Two wizard modes share that path. Quick Start asks four questions, the project's name, what it does, who it is for, and whether people will sign in. Everything else takes the default, Next.js, FastAPI, PostgreSQL, Render, and a runner that pushes its own deploys. A Working Mode section in the project's CLAUDE.md keeps every message in plain language, turns blockers into short questions with options, and takes answers from chat. Custom asks a developer about stack, auth, env group, and who pushes. Both render from one config into the same templates, so there is one test surface.
07The first full build
On September 18 I ran the whole thing end to end on a Quick Start project called factory-notes. The non-interactive wizard created the repo and provisioned Render. /spec create ran headless with the answers supplied, and a shell loop drove /orchestrate as a stand-in for /goal. The brief had four requirements.
Brief 001 reached 4-done in one runner turn. The runner pushed twice under its own push policy. The backend-worker's 11 tests ran against the real database. The frontend, a corkboard of notes, was verified by a local screenshot and again on the live URL. Both deploys went live and the health checks returned 200. Then provision.py --destroy tore it all down.
The run surfaced three things worth fixing. Headless runs skip hooks and project settings unless the project is marked trusted, which interactive users handle through the trust dialog. Homebrew builds Node from source on macOS 13, so /start now installs the official prebuilt tarball into ~/.local. And the frontend-worker left its test notes on the live board. That was acceptable for a smoke test and earns a cleanup step in real projects.
08Why this shape
- The folder is the status. Anyone can read the board with
ls. A card in 3-blocked is a person's to-do list, and a card in 4-done carries evidence for each requirement in its Outcome section. - One writer, many workers. Workers get fresh context, a narrow file list, and a report to fill in. Everything they learn flows back through the one session that owns the brief.
- Infrastructure first. The database and services exist before the first feature. The backend tests against the database it will run on, and the frontend talks to a deployed API from its first screen.
- Two terminals. Parking blockers keeps unrelated work moving. A separate 3-blocked folder keeps a stuck brief from ever being reported as a success.
- Hooks for the rules that matter. Documentation, trajectory, and the workspace pin are enforced by the harness, so they hold on turn fifteen as firmly as on turn one.
Consul belongs to a small family of Claude Code tools. Overtone gets instructions into a session by voice, Relay hands tasks to Claude Code from a to-do list, and Consul takes a description all the way to a running app. The lesson that carried across all three is that agent work becomes trustworthy when its state lives somewhere a person can look and its rules live somewhere the model cannot forget.
Consul is open source under MIT at github.com/nbdesai1992/consul. It needs a Mac, Claude Code, a GitHub account, and a Render account with a payment method. First-time setup takes about twenty minutes, and a small live project costs about $20 a month on Render. Companion pieces: Overtone, which puts your voice into Claude Code, and Relay, which delegates to Claude Code from your to-do list.
# once per machine: clone, open Claude Code, and let /start do the rest
git clone https://github.com/nbdesai1992/consul.git ~/consul
cd ~/consul && claude
/start
# then, inside the new project
/spec create "an invoice tracker for freelancers"
/goal <the line the spec skill hands you>
/status