team-bootstrap — AI delivery framework

Open-source role-based delivery framework for Claude Code: a proof-of-delivery layer over spec-driven development, with pipelines, an independent verifier and hook-enforced guardrails.

Role
Creator & maintainer
Stack
Claude CodeShellOpenTelemetryMCPSpec Kit
A delivery pipeline — spec, plan, build, independent verify, ship — under hook-enforced guardrails, OpenTelemetry tracing and an eval harness.

team-bootstrap is my open-source framework for making coding agents deliver reliably instead of ad hoc. It runs a software-engineering task through Product, Architecture, Implementation, Review and Release roles inside Claude Code — with structured handoffs, validation and observability.

Source: github.com/polischuks/team-bootstrap · MIT · v4.

What it is, precisely

Spec-driven development (SDD) tools take an idea to code and stop there — GitHub Spec Kit is explicit that it produces the artefacts and does not verify that the implementation satisfies the specification. team-bootstrap is that missing half: a proof-of-delivery layer over SDD. Its subject is closure-time verification — which roles a change earned, whether they actually ran as independent minds, and whether a batch may be called done. It runs the pre-implementation flow through Spec Kit’s own commands rather than replacing them.

It is deliberately not a harness. Claude Code is the harness; team-bootstrap is a policy layer on top of one, which is why every lever it needs is requested through the host’s hook API rather than asked for in prose.

Design

It is single-thread by default: roles are output styles activated for distinct phases of one Claude session, sharing a run document as a blackboard. Subagents are dispatched only for context isolation — research, security audit, parallel reviews — never to delegate a decision. That follows Cognition’s “Don’t Build Multi-Agents” principle: share context broadly, use subagents narrowly. The multi-role pipelines stay available for work that needs formal phase gates and an audit trail.

The design is pinned by a versioned constitution of invariants (P1–P12) that every milestone must respect — among them: harness-enforced policy with the LLM out of the security loop (P3), irreversibility gated behind approval (P5), typed schema-validated handoffs (P4), verification by red→green evidence rather than assertion (P9), and verification that is cumulative and fail-closed (P10).

Launch pipelines

The task is matched to a pipeline, not run through one fixed orchestration:

Pipeline What runs When
single-thread one session, three phases: plan → implement → verify most engineering tasks (the default)
mvp 7 roles: product-ba → delivery-manager → cto-architect → backend → frontend → qa → release-docs internal, low-risk, quick iterations
full 20 roles: adds discovery, formal product/business/test design, specialized reviewers, product & growth marketing (GTM), a release manager, stakeholder comms and docs production, customer-facing, compliance-sensitive
role one targeted role, by name when only a single phase is needed
audit 15 roles, read-only technical/operational readiness → remediation backlog
l2p 6 roles, evidence-disciplined landing↔platform↔docs gaps → ICE-ranked backlog
audit-dd 6 due-diligence roles, read-only investment / M&A / board review → investor-grade memo

Go-to-market is modelled as roles, not a separate pipeline: product-marketer (ICP, positioning, pricing, launch sequencing), growth-marketer (channels, content engine, AI-search posture) and partnerships-lead fold into full — and are available single-thread — triggered by a new product, a new ICP, repositioning, a pricing change or a new GTM motion.

/deliver chains the whole thing into one entry point for a spec-driven milestone. Phase A runs autonomously — constitution → specify → clarify → plan → tasks → analyze — and stops on a hard blocker or a CRITICAL inconsistency. Phase B decomposes tasks into batches and fires them one at a time through the chosen pipeline, waiting for confirmation between batches; subagents commit locally and nothing is pushed without explicit authorization. With no tier pinned it sizes the run itself, reading tasks.md/plan.md to derive a per-work-stream role plan. Interrupted runs resume from the last completed handoff; runs replay from their trace for prompt-regression evals.

How delivery is verified

Trust comes from enforcement, not prose. The policy is carried by 46 fail-closed check scripts and by Claude Code hooks (PreToolUse, Stop/SubagentStop, SessionStart, UserPromptSubmit, PreCompact) — so a rule is enforced by the host, not by hoping the agent remembers it. Handoffs between roles are typed and schema-validated; an independent reviewer, an integration verifier and a regression guardian check the work rather than letting the builder grade its own homework. Every run is traced with OpenTelemetry, and an eval harness measures whether a change to the framework actually makes agents deliver better.

The run document

Every run produces a single markdown run document: run metadata, a section per role with its output and a YAML handoff, and a final verdict — a release decision or a blocker list. It is the canonical audit record and the input to the grading evals. “Blocked” is always preferred over a false “complete”.