Crafting AI Prompts Framework

The CRAFT Framework: Orchestrating Agentic Flow

Last updated: Jul 23, 2026

Generative AI changes how software creates value. Traditional agile frameworks rely on heavy, synchronous human alignment to push work through a pipeline; CRAFT is built for the era of AI. It is a lightweight, asynchronous method for managing context, setting firm risk boundaries, and orchestrating AI agents to build high-quality software fast.

The Core Philosophy

Generative AI has enormous execution speed but no built-in direction or business context. CRAFT rests on one premise: humans define the boundaries; AI accelerates the execution. Rather than timeboxed sprints that cap velocity, CRAFT teams work in continuous flow. Asynchronous communication, automated state tracking, and disciplined instruction management remove the bottlenecks of legacy development.

This flow rests on three foundational pillars. Note: Security and Sustainability are not separate pillars; they are constraints embedded in all three.

  • People: Human intent cannot be automated. People move from writing boilerplate to defining business value, setting risk boundaries, and orchestrating AI behaviour. Adoption and team health are actively managed to prevent burnout at speed.
  • Processes: Continuous, asynchronous workflows replace synchronous ceremonies. The process feeds the AI exact context just in time and verifies output without stalling developer momentum.
  • Technology: The execution engine: the AI models, MCP servers, CI/CD pipelines, Platform Engineering paved roads, observability, and the live cockpit that make the framework’s speed possible.

The Always-On Principle. Traditional teams create value only while humans are at their desks. CRAFT rejects that limit: AI agents have no working hours, and run around the clock when properly instructed. Treat every off-hours window (evenings, weekends, holidays) as execution capacity. The human role shifts from constant supervision to asynchronous orchestration: prepare the work, dispatch the agent, step away, then return to review results and clear HITL checkpoints. (More in the Always-On Execution note below.)

System Over Output. Traditional teams are measured on output: features shipped, tickets closed, pull requests merged. CRAFT changes what the team is accountable for. Its real job is to build and continuously harden the system that produces the output, not to produce it by hand. That system is the harness: the instructions, context, tools, guardrails, tests, and observability that turn a raw model into a reliable engineer. Build the harness well and quality and speed improve on their own, without adding people. Let it drift, through context rot, schema misalignment, or instruction bloat, and the output drifts with it, however hard anyone works. For anyone arriving from a traditional engineering background, this is the mental shift: you are no longer judged mainly on what you shipped this week, but on whether the system you steward is smarter, faster, and more reliable than it was last Micro-Cycle.

The Leadership Contract

CRAFT teams move fast because they have autonomy, and that autonomy has to be granted and protected by leadership. The research is blunt: 70–88% of digital transformations fail on leadership misalignment, not technology (BCG; Bain 2024). Gartner projects that organisations prioritising executive AI literacy will see 20% higher financial performance by 2027. CRAFT handles this with a short, two-way contract.

What leadership must provide:

  • Trust asynchronous execution. The State of the Craft cockpit is your status report, and it is always live. All requests flow through the Flow Orchestrator, never directly to engineers.
  • Protect the team from organisational antibodies. SAFe ceremonies, cross-department steering committees, and “quick check-in” meetings will try to re-impose synchronous overhead. The executive sponsor shields the team from them.
  • Shift the budget model. CRAFT teams need fewer people but more infrastructure: model API budget, MCP access, premium tooling. Approve these as operating spend, not project-by-project procurement.
  • Resist the headcount reflex. When you need more velocity, invest in better tooling and Core Craft quality, not more people. Adding headcount to a CRAFT team rarely helps and often hurts.

What leadership receives in return:

  • Six numbers that matter (detailed in the Measuring Success section): Blueprint Cycle Time, Gate 1 First-Pass Rate, Effort Weight Accuracy, HITL Frequency Rate, API Cost per Blueprint, and Calibration Health Score. They answer three questions: How fast are we shipping? Can we trust our forecasts? Is this sustainable?
  • A Monthly Business Review (30 min) replacing steering committees: metrics, outcomes delivered, and strategic redirect. No slide decks, no demos of half-finished work.
  • A Quarterly Strategy Alignment (90 min) replacing PI Planning: leadership sets direction (“what business problems to solve”), the team decides how.
  • A Quarterly Business Value Report connecting CRAFT metrics to business outcomes: value delivered, efficiency gains, risk posture, sustainability health, and forward forecast.

Leadership sets direction and removes barriers; the team decides how to execute and delivers measurable outcomes. Anything outside these channels (ad-hoc status requests, direct messages to FDEs, surprise steering committees) is an anti-pattern that degrades the system it means to help.

The CRAFT Ecosystem: Core Roles & Boundaries

In an agentic workflow, people are not managed as resources; they are empowered as the “expert in the lead.” By handing repetitive execution to the AI, team members step fully into their domain expertise. A CRAFT team is a small, cross-functional unit of specialized focus areas.

Team shape and size. Because Generative AI is a force multiplier, CRAFT teams stay exceptionally lean, typically 2 to 4 people and often fewer. The building is done by one or more Forward Deployed Engineers, paired with a Harness Engineer who owns independent verification and the health of the system. A typical shape is one Harness Engineer to one to three FDEs, a ratio that works because verification is automated: one person can be the independent verifier and harness owner for several builders when Gate 1 does the heavy lifting. As agents take on more of the execution and verification load, expect this number to shrink, not grow; adding people when you need velocity is an anti-pattern (see the Leadership Contract).

Staffing for stand-by duty is the one legitimate reason to add people. CRAFT’s always-on model means agents, and the incidents they can trigger, run off-hours, so any team with an on-call rotation needs enough people to cover it without burning anyone out. A two-person team on a 24/7 rotation is a sustainability failure the Calibration exists to catch. Note the reason this is allowed: it is coverage-driven, not velocity-driven. Adding people to ship faster is still an anti-pattern; adding people, or pooling the rotation across several teams, so the stand-by load stays humane is fine. If a Harness Engineer becomes a review bottleneck as throughput grows, the first fix is more automation in the harness, not another Harness Engineer.

Which roles are shared, which are embedded. Not every role maps one-to-one to a single team. The Product Strategist and Flow Orchestrator are genuinely multi-team: one Strategist can sequence and prioritize across several teams, and one Orchestrator serves one to several teams as the connective tissue between them, the organisation, and the wider platform ecosystem. The Harness Engineer leans embedded instead, because independent verification and daily harness tuning depend on proximity to the team’s Blueprints; only the deepest specialist layers (enterprise AppSec, shared eval and platform infrastructure) sit above the team as a shared capability. In large enterprises those deep experts, such as AppSec engineers and central Platform Engineering teams, act as a Shared Service across many CRAFT teams, because with AI they can build rules, components, and infrastructure fast enough to support the whole portfolio. This is a structural principle, not cost-cutting: teams stay lean because shared expertise scales across the portfolio rather than being duplicated in every team.

Small-team role flexibility. In small or early-stage teams, one person may hold more than one role, most often combining the FDE and Harness Engineer. This is allowed. What is never allowed, at any size, is reducing the harness and QA accountability itself. A two-person team applies the same adversarial verification, the same Gate 1 automation, and the same HITL checkpoints as a twenty-person team. The team is smaller; the bar is identical. This works because verification is mostly automated (Gate 1 runs the same pipeline whether one person or ten triggered it), so team size does not set verification rigor; harness configuration does. What does take discipline is the human judgment layer: when one person builds and reviews, they must force a context switch. Finish the build. Close the builder mindset. Reopen the work as the Harness Engineer with one question: “How would I break this, and what did the agent get wrong?” Blending both into one unchecked pass defeats the boundary. In regulated enterprises, separation of duties may be legally required regardless of team size; check your compliance requirements before combining roles.

Looking ahead: role boundaries in a more agentic world. Today CRAFT enforces role separation through human discipline: the Strategist decides what, the FDE builds how, the Harness Engineer verifies whether it is safe and optimal. As agents get more capable, this discipline-based model will be tested. Agents already draft Blueprints from call transcripts (see Agent-Initiated Blueprints), run their own risk analyses, and execute code, blurring who initiates what. Before long, a single agent system may cover work that today spans the Strategist, FDE, and Harness Engineer at once. CRAFT answers this with one principle: the origin of work may be automated, but the verification of work must stay structurally separated. Concretely: (1) Core Craft permissions must enforce which agents can propose versus approve; the agent that drafts a Blueprint must never be the instance that approves its risk analysis or signs off Gate 2. (2) As autonomous workflows grow, the Flow Orchestrator keeps an Agent Capability Registry, a living inventory of which agents do what, with boundaries that stop any single agent from spanning the full propose–build–verify chain without a human or independent-agent checkpoint. (3) HITL thresholds become more important, not less, as agents improve. The temptation to raise them (“it’s right 95% of the time, skip the review”) must be resisted until governance agents can provide equivalent adversarial verification. The gates and checkpoints are not overhead to optimise away; they are the structural guarantee that speed does not outrun accountability.

Roles within CRAFT

The CRAFT Engine: Executing Agentic Flow

CRAFT drops traditional sprint ceremonies. The workflow runs as continuous, asynchronous phases called Micro-Cycles (typically 3 to 7 days), built to keep the pipeline fed and the AI executing. Each phase has a clear owner, a concrete output, and an explicit handoff, so everyone knows what to do and when.

What is a Micro-Cycle? A Micro-Cycle is a team-level cadence period, not a per-Blueprint sprint. Blueprints start and finish continuously within and across Micro-Cycles; no sprint boundary freezes the backlog or batches work artificially. The boundary exists for one thing only: to trigger the CRAFT Evolution meeting, the team’s heartbeat for system hardening. The team sets the duration in its Team Agreement and shows it on the cockpit. It should track the rhythm at which meaningful agent learnings accumulate, not a fixed calendar slot. This is a team process decision, not an AI agent instruction, so it does not belong in the Core Craft.

Note: this section covers the workflow phases and how the team operates. The artifacts that power them, the Blueprint, Core Craft, and Local Craft, are defined in the Knowledge Threads section below. Read both together for the full picture.

Note: all meeting timeboxes here are baselines. Because CRAFT runs fast and evolves continuously, teams should adjust these durations when different constraints serve them better.

1. Blueprint CRAFTing & Agentic Forecasting (Async & Ongoing)

A Blueprint is the single source of truth for a unit of work. It is not a traditional requirements document; it is a machine-readable, AI-executable specification that holds everything the agent needs to act and everything the team needs to verify the outcome. Blueprints are created continuously; there is no “sprint planning” gate. As soon as a Blueprint is fully formed and forecasted, an FDE can start executing it.

Scope. A Blueprint covers the complete scope of a unit of work, from UI changes and API logic down to database schema updates. Its boundary is a coherent business outcome, not a technical layer. For very large features that genuinely need multiple sequential Blueprints (for example, the data model must ship before the API can be built), group them under a Blueprint Epic: a lightweight parent record that links dependent Blueprints in order and makes the sequence explicit in the backlog. An Epic is not a planning ceremony, just a named sequence.

Backlog vs. the cockpit. How the team manages its Blueprint backlog is its own choice. Some teams keep Blueprints as markdown files in the repo (managed via a backlog.md); others use a board tool. The framework mandates no tool; pick whatever the team will actually use and maintain. What matters is one source of truth for the Blueprint queue. The State of the Craft cockpit is not the backlog; it shows only real-time operational state (what is executing, blocked, or staged for verification). Backlog management and prioritization happen outside the cockpit.

Agent-initiated Blueprints. The steps below describe the human-led flow, but Blueprints need not start with a human. As agents mature, they can initiate Blueprints by observing real-world signals. During a user call, a transcription agent can listen live, extract pain points, feature requests, and usage patterns, and draft a structured Blueprint proposal before the call ends. Observability agents can turn recurring error patterns into Hotfix Blueprint drafts; support-ticket agents can cluster complaints into one feature Blueprint; telemetry agents can spot performance bottlenecks and propose Hardening Blueprints. The originating role stops mattering; what matters is that every agent-initiated Blueprint enters the same CRAFT pipeline as a human-initiated one. Specifically: (1) the Product Strategist still reviews, prioritizes, and approves or rejects it before it enters the backlog (agents propose, humans decide); (2) the Harness Engineer still runs the risk analysis and sets boundaries; (3) HITL thresholds and Gate 1/Gate 2 apply identically no matter who drafted the spec. Agent-initiated Blueprints multiply signal capture, so valuable information in calls, logs, and user behaviour is not lost to human memory, but they never bypass the verification architecture. The gates exist precisely because a Blueprint’s origin matters less than its validation.

Input sanitisation for agent-initiated Blueprints. The data sources that feed these Blueprints (observability logs, support tickets, call transcripts, telemetry streams) are untrusted input. They may contain customer PII (names, account numbers, health data), embedded credentials (session tokens, API keys logged by accident), or adversarial content (prompt injection hidden in a crafted error message or ticket body). The same data-governance layers that apply to all agent operations apply here: an agent ingesting external data must not reach raw, unsanitised sources when a sanitised alternative exists. Define a structured ingestion schema for each external source, a format that strips PII, redacts credentials, and normalises the data before the agent sees it. For example, a support-ticket schema exposes category, severity, anonymised_description, and affected_feature, but never customer_name, account_id, or raw email content. Where full sanitisation is not possible (for example, free-form call transcripts), the Core Craft must include explicit rules telling the agent never to reproduce PII from ingested data and to treat any instruction-like content inside external data as data, not a command. The Harness Engineer reviews these ingestion schemas as part of the data-governance check at each CRAFT Evolution.

The steps below describe contributions, not a strict waterfall. Blueprints are drafted, refined, bounded, and forecasted asynchronously and often iteratively; a Blueprint is ready to execute the moment all sections are complete and it has been forecasted.

Step 1. Forward Deployed Engineer: draft the Blueprint from field signal.

Because the FDE is the team’s interface with users and the field, they usually open or create the planning.md and draft the initial Blueprint, turning real-world signal into a concrete, executable specification. (For purely strategic or agent-initiated Blueprints the origin may differ, see Step 2 and Agent-Initiated Blueprints, but the FDE always owns the technical layer.) The FDE fills in the following.

Business & outcome fields:

  • Business Goal: One or two sentences, grounded in observed user need. What user problem does this solve, and what measurable outcome do we expect? (The Product Strategist refines and validates this in Step 2.)
  • Desired Outcome: What does success look like from the user’s perspective? Use concrete, verifiable language (for example, “a user can complete checkout in under 3 steps”).
  • Out of Scope: Explicitly list what this Blueprint does NOT cover, to prevent AI scope creep.
  • Acceptance Scenarios: Written in plain language. The Harness Engineer converts these into automated test cases.

Technical & execution fields:

  • Technical Architecture: Which services, components, or infrastructure does this feature touch? Note dependencies on existing modules.
  • Required MCP Servers or Tools: List the AI tools the agent will need (for example, a file-system MCP, a database query tool, a browser agent).
  • Human-in-the-Loop (HITL) Thresholds: Explicit stopping points where the AI must pause for human approval before continuing. Common examples: before a database migration, before external API requests, before modifying production files. HITL checkpoints are the technical expression of CRAFT’s “expert in the lead” principle: the precise moments where human expertise overrides automated momentum and reclaims control of execution.
  • Local Craft Reference: Note relevant rules from dev-instructions.md or known agent hallucinations that apply to this feature.
  • Context Loading Instructions: Define which files, modules, and documentation the agent must load before execution, and which it must not load. Agents with full-repo access often pull irrelevant context that degrades output (“context rot”: as the window grows, attention quality drops). Curated context loading is one of the highest-leverage improvements to agent output quality. Example: “Load src/payments/ and docs/payment-api.md. Do NOT load src/legacy/ or tests/e2e/.” Include any live data lookups (schema snapshots, API specs) the agent needs but that should not live permanently in the Core Craft.
  • Agent Execution Constraints (Loop Detection): Hard guardrails for this Blueprint’s execution: maximum tool calls before an automatic HITL halt (25–50 for typical feature Blueprints), maximum wall-clock time before auto-halt (30–60 minutes), and a maximum token-spend threshold that triggers an automatic pause and a cost-overrun alert to the Flow Orchestrator. These override the global Core Craft limits for this Blueprint only; with no Blueprint-specific values, the agent runs on the global defaults. Runaway agents are both a budget risk and a trust risk: teams that have watched an agent silently spend 10× the expected cost on one Blueprint rarely trust agentic execution again without discipline here.
  • Depends On (Optional): If this Blueprint cannot run until another is deployed, list its ID or title. The Blueprint is marked blocked until its dependency completes, which stops FDEs building on infrastructure that does not yet exist in production.

Off-hours execution readiness. Evaluate every Blueprint for off-hours executability during Blueprinting. The FDE tags one of three levels: (a) fully autonomous, the agent can finish with no HITL interaction, fit for overnight or weekend dispatch; (b) low-touch, 1–2 HITL checkpoints that can be batched and cleared in one review the next morning; (c) high-touch, needs continuous human steering during working hours. Levels (a) and (b) are prime end-of-session dispatch candidates. Teams that only ever produce (c) Blueprints should revisit their Blueprinting discipline; vague specs and missing context are the main reasons an agent cannot run autonomously. Blueprints designated as (a) or (b) can be automatically queued and picked up by agents in a pipeline that aligns with the Product Strategist's needs (and release strategy).

Step 2. Product Strategist: prioritize, sequence, and optimize.

The Strategist does not usually author the draft; their accountability is to decide whether, when, and in what order the Blueprint is built, and to sharpen its business value. Working with the FDE, the Strategist:

  • Refines the Business Goal and Why, making sure the Blueprint targets a genuine, high-value user problem, not just the pain most visible to one embedded FDE.
  • Assigns priority using business value against the AI-calculated Effort Weight (from Step 4), and finalises the Blueprint’s place in the backlog.
  • Owns sequencing and pace: places the Blueprint in the right Release Wave and manages user absorption so the product is not flooded with changes.
  • Works with the FDE to tighten ambiguous specs before execution; a vague Desired Outcome or fuzzy scope is fixed here, not discovered at Gate 2.

For strategic Blueprints that come from product vision rather than the field, the Strategist writes the business What and Why and hands the draft to the FDE for the technical layer.

Step 3. Harness Engineer: set risk and verification boundaries.

Asynchronously, the Harness Engineer reviews the Blueprint and adds the following into planning.md:

  • Product Risk Analysis (PRA): A rapid assessment of what could go wrong. Rate each risk by likelihood and impact on this baseline matrix: Critical (high likelihood × high impact), mandatory HITL checkpoint and Harness Engineer sign-off before execution; High (high impact or high likelihood), HITL checkpoint strongly recommended plus a targeted automated scan; Medium, automated scan only; Low, log and monitor. Risk scores are recorded in the Blueprint and shown on the cockpit.
  • Threat Model: Identify attack surfaces the feature introduces (new API endpoint, file upload, user-data access) and define the mitigation that must be implemented.
  • Automated Test Constraints: Define the exact test coverage the pipeline must pass before deployment. These are non-negotiable gates, not suggestions.
  • Acceptance Criteria & Verification Harness: Translate the acceptance scenarios into binary pass/fail criteria that the pipeline and the Strategist can verify independently, and specify any golden-trajectory or adversarial coverage the Blueprint needs.

The FDE builds the tests into the software during execution (Step 1’s architecture made it testable); the Harness Engineer defines the independent bar those tests must clear.

Technical revision note: the FDE’s technical contributions (architecture, MCP servers, HITL thresholds, context loading, execution constraints) live in Step 1. If the Harness Engineer’s risk boundaries require a change to the technical approach, for example adding a HITL checkpoint the PRA flagged as mandatory, the FDE updates the relevant fields before execution.

Step 4. AI Agent: Agentic Forecasting.

Once the Blueprint is complete (the FDE’s draft and technical layer, the Strategist’s prioritization, and the Harness Engineer’s risk and verification boundaries all in place), an AI agent runs an automated analysis:

  • It reads the Blueprint against the global Core Craft (architecture constraints, coding standards, security rules).
  • It cross-references historical team throughput and codebase complexity metrics.
  • It outputs an Effort Weight, a normalized score (not story points) for relative execution complexity.
  • The Product Strategist uses the Effort Weight alongside business value to calculate ROI and place the Blueprint in the prioritized backlog. The highest-value, lowest-effort Blueprints execute first.

Blueprint Definition of Done

A Blueprint is only considered done when all of the following are true. This checklist can be embedded in planning.md so the AI agent can self-verify before marking the Blueprint complete:

  • Gate 1 passed: All automated acceptance criteria tests, SAST/DAST scans, and dependency checks passed without manual override.
  • Gate 2 passed: The Product Strategist has verified the feature against the business goal and desired outcome in staging.
  • Observation window closed cleanly: No regressions, error rate spikes, or security alerts surfaced during the post-deployment observation period.
  • Local Craft updated: Any AI hallucinations, HITL halts, or prompt failures encountered during execution are logged in the FDE’s dev-instructions.md.
  • Promotable learnings flagged: Any reusable architectural pattern, workflow, or technique worth sharing has been marked for the next Evolution’s Promotion Loop.
  • Off-hours execution tag assigned: The Blueprint’s FDE section explicitly tags the work as fully autonomous, low-touch, or high-touch for off-hours execution. This tag is set during Blueprinting and validated after completion; if the actual HITL interaction count deviated significantly from the tag, the FDE logs why in their Local Craft to improve future tagging accuracy.
  • Monitoring confirmed: Observability signals (logs, metrics, traces) for the shipped feature are active in the production environment.
  • Execution context snapshot logged: The model version, Core Craft version (commit hash), active MCP servers, active skills, and Blueprint version at execution start are recorded alongside the Blueprint’s completion record. This lightweight provenance trail enables the team to answer “what instructions and tooling were active when this Blueprint was built?” without manual Git archaeology. When debugging unexpected agent behaviour or evaluating the impact of a Core Craft change, the snapshot makes it possible to compare execution conditions across Blueprints. The Flow Orchestrator owns the tooling setup that captures this automatically; where there is no automated capture, the FDE logs it manually at execution start. This snapshot is also the basis for attribution (answering “which model, instructions, and tools produced this code?”), a governance capability increasingly expected as AI-generated code reaches production.
  • Execution trace posted to the issue: A structured comment is posted to the originating issue or ticket (see Execution Traceability below) so the Harness Engineer and FDE can reconstruct exactly what happened without opening the agent session or digging through Git history.

Execution Traceability: The Issue as the Audit Log

Every CRAFT Blueprint originates from an issue or ticket. That issue is already the natural place to record what happened during execution — not just that it was done, but how it was done. The prompt that drove the agent, the key decisions it made, the output it produced, and a link to the resulting PR all belong in the issue, posted as a structured comment when the agent completes its run.

This matters most for (A) Agentic Blueprints, where no human is present during execution. Without a trace in the issue, the Harness Engineer has no reliable way to review what the agent did, the FDE has no starting point when something goes wrong, and the team cannot answer the most basic governance question: what exactly produced this code? But it is equally valuable for (B) Balanced runs — HITL approval decisions should be recorded in the issue too, so there is a clear log of what the human saw and chose to approve at each checkpoint.

What the trace comment must contain:

  • Execution mode and label: Which label triggered this run (A / B / C) and the dispatch mode.
  • Model and harness version: The model name and version, Core Craft commit hash, active MCP servers, and active skills at the time of execution. MCPs define what external tools the agent could reach; skills define what domain-specific instructions and capabilities it carried. Both directly affect what the agent did and how — if either changes between runs, the output may differ. This is the same data as the execution context snapshot in the DoD, surfaced here in the issue so it is immediately visible without opening a separate record.
  • Key prompt(s) used: The primary instruction or Blueprint excerpt that drove the agent. Not the full conversation transcript, but enough to understand what the agent was told to do and what constraints it was given. For (A) Agentic runs this is especially critical — it is the only human-readable record of what drove the autonomous execution.
  • Execution summary: A brief, structured summary of what the agent did: steps taken, tools called, decisions made at branch points, and any self-corrections. Most agents can generate this as a closing summary if instructed to do so in the Core Craft.
  • Output reference: A direct link to the PR or merge request, the specific commit hash, and (if applicable) the Gate 1 run result. This closes the loop between the issue and the code it produced.
  • Deviations and anomalies: Anything unexpected: a prompt that had to be adjusted mid-session, a tool call that failed and was retried, an output that diverged from the acceptance criteria in a minor way. These are exactly what the Harness Engineer needs to spot patterns and tighten the harness.
  • HITL decisions (B only): For each checkpoint where a human approved continuation: who approved, at what step, and what they saw before approving. This makes the human judgment part of the audit log, not just the agent’s actions.

Who posts it and when: For (A) Agentic runs, the agent posts this comment automatically at the end of its session, before it opens the PR — this should be a Core Craft instruction so every agent does it consistently. For (B) Balanced runs, the agent posts the execution summary and the FDE appends HITL decisions. For (C) Crafted runs, the FDE posts a brief note of the approach taken if the implementation involved a non-obvious prompt or technique — this builds institutional knowledge even for human-led work.

Why the issue, not just the PR: Pull and merge requests are excellent for reviewing the diff, but they are poor audit logs for agent behaviour. Comments disappear into review threads; the PR is closed and archived once merged. The issue stays open as a permanent record linked to the code change, searchable by the Harness Engineer when a pattern surfaces months later. The issue is also where the Execution Mode label lives — keeping the trace there means all the information about a Blueprint (what was asked, what ran, what was produced) lives in one place.

CRAFT Pipeline Labels: Automating the Workflow

Labels are the connective tissue between human intent and automated pipeline execution in CRAFT. When applied consistently in your issue tracker — GitHub Issues, Jira, Azure DevOps, Linear, or any equivalent — they tell the pipeline exactly what to do, how much human oversight is required, and whether the work can run unattended overnight. Without a shared label vocabulary, every team member makes different assumptions about who should act next and whether the agent can proceed autonomously. With it, a single label on an issue triggers the right branch, the right agent, and the right review flow automatically.

CRAFT labels fall into five categories. Every issue should carry at minimum one label from each of the first two categories (Execution Mode and Blueprint Type) before it enters the backlog. Risk, State, and Dispatch labels are added during Blueprint CRAFTing and evolve as the work progresses through the pipeline.

Category 1 - Execution Mode (required on every issue)

Determines how much autonomous pipeline execution is triggered. Assign exactly one per issue. This is the label the automation pipeline reads first.

Label NameHuman / AIDescription
(A) — AgenticAI onlyFully autonomous pipeline execution. The agent creates a branch, implements the work, opens a pull or merge request, and runs Gate 1 checks without any human involvement during execution. A human reviews the output only at the Gate 1 PR stage. Use for well-specified, low-risk Blueprints with clear acceptance criteria.
(B) — BalancedAI + HumanAgent-driven execution with mandatory HITL checkpoint(s). The pipeline starts the agent automatically, but halts at defined checkpoints for human approval before continuing. The correct default for most feature work: agent speed with human oversight at the moments that matter. Use when the Blueprint has medium risk, external API calls, database changes, or any step that requires judgment before proceeding.
(C) — CraftedHuman onlyEngineer-led; AI is a copilot, not the driver. No autonomous pipeline execution is triggered. The engineer owns design decisions and implementation direction. Use for architectural work, novel security boundaries, regulatory compliance requirements, or any context where autonomous execution risk outweighs its speed benefit.

Category 2 - Blueprint Type (required on every issue)

Describes the nature of the work. Used to route issues to the correct pipeline variant and to calculate meaningful throughput metrics across issue types.

Label NameHuman / AIDescription
type: featureA or BA new user-facing capability or business outcome. Should always have a Blueprint with a defined business goal, desired outcome, and acceptance scenarios.
type: bugA or BA defect in existing functionality. Include repro steps and expected vs. actual behaviour in the issue body. Well-described bugs in bounded areas are good candidates for (A) Agentic execution.
type: hotfixB or CAn urgent production defect requiring immediate resolution. Hotfixes almost always warrant at least (B) Balanced due to time pressure and production risk. The Harness Engineer must be looped in immediately; standard Gate 1 automation still applies but may run in parallel with human triage.
type: hardeningA or BA Hardening Blueprint: zero new user-facing functionality, dedicated to refactoring, test-coverage improvement, observability gaps, or technical debt reduction identified at the CRAFT Evolution. Count toward team throughput like any other Blueprint.
type: epicHuman (C)A Blueprint Epic: a parent issue that groups multiple dependent Blueprints in sequence. Not directly executed; serves as a tracker. Child Blueprints carry their own Execution Mode and Type labels.
type: craft-evolutionHuman (C)A Core Craft or harness improvement identified at the CRAFT Evolution: a new rule, guardrail update, or system-level change that has no direct user-facing output but strengthens the agent harness.

Category 3 - Risk Level (from the Harness Engineer’s PRA)

Set by the Harness Engineer during Blueprint CRAFTing based on the Product Risk Analysis. Determines the number of mandatory HITL checkpoints and whether additional sign-off is required before Gate 1.

Label NameHuman / AIDescription
risk: lowAI primaryNo significant security, data, or system-stability risk identified in the PRA. No special HITL checkpoints beyond the standard Gate 1 human review. Compatible with (A) Agentic execution and off-hours dispatch.
risk: mediumAI + HumanOne or more HITL checkpoints required (e.g., before external API calls, before schema changes). Agent can execute between checkpoints autonomously, but must halt for approval at defined steps. Compatible with (B) Balanced; may still be dispatched off-hours if HITL notifications reach the responsible human promptly.
risk: highHuman gateMultiple HITL checkpoints; Harness Engineer sign-off required before the PR can be merged. Typically involves production infrastructure changes, security boundary modifications, data migrations, or regulatory compliance scope. Requires (B) Balanced or (C) Crafted. Do not dispatch off-hours without explicit team agreement.

Category 4 - Dispatch Mode (set during Blueprint CRAFTing, validated at Definition of Done)

Controls whether the agent can run unattended (e.g., overnight, over a weekend). Directly feeds the Always-On Execution principle. Set in the Blueprint alongside the Execution Mode label; validate accuracy after completion and log any deviation in the Local Craft.

Label NameHuman / AIDescription
dispatch: off-hoursAI primarySafe to dispatch at end-of-session for overnight or weekend execution. The Blueprint is sufficiently specified and the risk level is low enough that the agent can run to completion (or a known checkpoint) without active FDE supervision. Combine with (A) Agentic + risk: low for maximum unattended throughput.
dispatch: supervisedAI + HumanCan run off-hours only if the responsible FDE or Harness Engineer has HITL notifications configured and can respond promptly to checkpoint halts. The agent will pause at HITL points and wait; if no response comes within a defined timeout (set in the Team Agreement), it should stop safely and log the halt.
dispatch: active-onlyHuman gateMust only be executed while an FDE is actively present and monitoring. Do not dispatch off-hours. Typically paired with risk: high, (B) Balanced, or (C) Crafted. Used for production infrastructure changes, data migrations, security boundary work, or any execution where an immediate human response to an unexpected agent action is required.

Category 5 - Pipeline State (managed automatically by the CI/CD pipeline)

State labels are applied and removed automatically by the pipeline as a Blueprint moves through the CRAFT Engine. They reflect real-time status on the cockpit without anyone manually updating a board column. Teams should not set these manually; they are owned by automation.

Label NameHuman / AIDescription
state: blueprint-readyAuto (pipeline)The Blueprint is fully specified (all required fields complete, Execution Mode and Risk labels set, Effort Weight calculated) and ready for an FDE to pick up and begin execution. Applied automatically once Blueprint CRAFTing is complete.
state: in-executionAuto (pipeline)The agent is actively working on this Blueprint. Applied automatically when a branch is created and the agent session starts. Visible on the cockpit to the whole team.
state: hitl-waitingHuman actionThe agent has reached a HITL checkpoint and is paused, waiting for human approval to continue. Applied automatically when the agent halts at a defined checkpoint. The FDE or Harness Engineer must respond; a lack of response within the Team Agreement timeout should trigger a notification escalation.
state: gate-1-blockedFDE fixesGate 1 automated checks failed (tests, SAST/DAST scans, or dependency alerts). The FDE is responsible for fixing the failures and re-pushing. The Harness Engineer monitors Gate 1 block rates; a rising trend is a Core Craft signal, not a people problem.
state: gate-2-pendingStrategistGate 1 passed; the Blueprint is staged and awaiting the Product Strategist’s asynchronous business verification (Gate 2). The FDE moves on to the next Blueprint; they do not wait for Gate 2 completion.
state: blockedFlow SyncExecution cannot proceed due to an external dependency, missing context, or unresolved decision. Triggers a Flow Sync within 30 minutes. The Flow Orchestrator is responsible for resolving or escalating. A Blueprint should never stay in this state for more than one working day without a documented escalation.

How the Labels Work Together: A Practical Example

A well-labelled CRAFT issue gives the pipeline, the team, and any automated triage agent a complete picture at a glance. For example, an issue labelled (A) — Agentic type: bug risk: low dispatch: off-hours tells the pipeline: create a branch immediately, dispatch the agent tonight, let it run to completion and open a PR — no human needed until the Harness Engineer reviews Gate 1 tomorrow morning. Contrast that with (B) — Balanced type: feature risk: medium dispatch: supervised, which tells the pipeline: start the agent, but stop at the schema migration step, notify the FDE, wait for approval, then continue.

The label combination is the Blueprint’s execution fingerprint. Standardising it across the team is what makes CRAFT consistent: any team member, at any time, can look at an open issue and know exactly how it will be handled — and any automation agent can read the same labels and act accordingly without ambiguity.

2. State of the Craft (Continuous Hybrid Cockpit)

The State of the Craft cockpit is the team's single, always-live source of truth. It replaces the daily standup entirely. No one asks “what are you working on?”; the cockpit already knows. Every team member is responsible for keeping their assigned work accurate on the cockpit in real time.

Tooling note: CRAFT does not prescribe a specific cockpit tool. Choose whatever your team will genuinely maintain: a GitHub Projects board with CI/CD webhook automation, Linear, Jira, or any equivalent. The only requirement is that the tool supports automated state transitions (triggered by pipeline events), real-time filtering by Blueprint status, and visibility of agent activity and pipeline health signals. The cockpit is a live operational view, not a project management archive.

What the Cockpit Tracks (Automated):

  • Blueprint Status: Each Blueprint moves through active states on the cockpit: In Progress → Verification → Staged → Done. A Blueprint only appears on the cockpit when an FDE picks it up from the backlog and marks it active; the backlog itself lives outside the cockpit. AI agents update these states automatically based on CI/CD pipeline events and repository activity.
  • Agent Activity Logs: What is each AI agent currently doing? Any loops, failures, or HITL halt points are surfaced immediately.
  • Pipeline Health: Build status, test pass rates, SAST/DAST scan results, and dependency vulnerability alerts.
  • Observability Signals: Key metrics, error rates, and traces from the staging environment so the team can spot regressions before they reach production.
  • Sustainability Metrics: Real-time API token burn rate, estimated compute cost, and carbon footprint for all active agents. Managed by the Flow Orchestrator, who sets hard budget thresholds and alerts.

What Each Role Does With the Cockpit:

  • Product Strategist: Monitors Blueprint progress and Agentic Forecast queue. Reprioritizes the backlog as new user signal arrives from the FDE. Flags Blueprints that are ready for their async verification review.
  • Harness Engineer: Monitors failed test gates and security scan alerts. If a critical vulnerability surfaces, they immediately update the Core Craft with a corrective guardrail and trigger a Flow Sync if it is blocking.
  • Forward Deployed Engineer: Updates their active Blueprint state. Logs any agent hallucinations or HITL halts they encounter in their Local Craft. If a blocker cannot be resolved within 15 minutes independently, they post it to the cockpit and initiate a Flow Sync.
  • Flow Orchestrator: Monitors the full team's API usage and compute footprint. Sets automated hard stops on any agent that exceeds defined token budgets. Identifies cross-team impediments and ensures the cockpit is always clean and actionable.

3. Flow Sync (Ad-Hoc Resolution - Max 15 mins)

A Flow Sync is not a meeting; it is a precision intervention. Its only purpose is to unblock the system as fast as possible. While humans resolve the blocker, AI agents keep executing everything else in the background. Calling a Flow Sync for anything that is not immediately blocking is an anti-pattern.

How to Initiate:

  • Any team member who hits a blocker they cannot resolve within 15 minutes posts it to the State of the Craft cockpit, tags it as a Flow Sync trigger, and pings only the team members whose input is required, not the whole team.
  • The Flow Sync must start within 30 minutes of being posted. If the required people are unavailable, the Flow Orchestrator unblocks the situation by escalating or making the decision.

Structure of a Flow Sync (15 minutes maximum):

  • Minutes 0–2, Context (FDE or blocker owner): State the problem in one sentence. What is blocked, and what has been tried? No background story, only the specific blocker.
  • Minutes 2–10, Resolution (relevant team members): Discuss only what is needed to decide. If it needs more than 10 minutes, the issue is too complex for a Flow Sync and should be re-Blueprinted.
  • Minutes 10–15, Call to Action (Flow Orchestrator): The Orchestrator closes the Sync with a concrete, named CTA: who does what, by when, written into the cockpit immediately. No action item left unowned.

What Happens After:

  • If the blocker is resolved, the cockpit is updated and execution resumes immediately.
  • If the issue turns out to be non-blocking, it moves to the backlog. The board stays clean; only active blockers remain visible.
  • If the Flow Sync reveals a systemic problem (for example, a repeated agent failure pattern), the Harness Engineer or FDE logs it in their Local Craft for the next CRAFT Evolution.

4. CRAFT Verification (Async)

The verification triad: who verifies what. CRAFT separates building from verifying, because the builder cannot be the sole judge of their own work. Three accountabilities divide the surface: the FDE builds the feature and its in-code tests (proving “does it work as I intended?”); the Harness Engineer independently verifies technical and system correctness through the automated harness and the human side of Gate 1 (proving “is it correct and safe, and does it clear the bar, whoever built it?”); and the Product Strategist verifies business value at Gate 2 (proving “did we solve the right problem?”). No one person owns all three. That separation keeps quality independent of individual optimism.

Verification is a two-gate process: an automated technical gate and an asynchronous human business gate. They run in sequence, never at the same time as active development, so the pipeline never becomes a bottleneck. The gates matter because an agent can convince itself its work is correct in its own local session while missing cross-service contract violations or environment-specific failures; CI verification (Gate 1), not the agent’s local confidence, is the real proof. The 2026 default loop is therefore orchestrate locally, verify asynchronously in CI. The FDE’s active build on a Blueprint is complete when the code is pushed to Gate 1; they do not take part in the Gate 2 business decision, which belongs solely to the Product Strategist. This is the asynchronous-first principle in practice: the FDE should not wait around for approval. They stay engaged, though: they fix Gate 1 failures promptly, watch the feature in staging, and remain accountable for its real-world outcomes through the production observation window. The build phase ends at the push; team accountability does not.

Gate 1: Automated Technical Verification (CI/CD Pipeline).

  • Triggered by (FDE or Agent): The FDE or their AI agent pushes the completed feature branch and opens a Pull Request. That is the only action needed to start Gate 1; every subsequent check runs on the PR automatically.
  • PR Review (AI-Assisted + Human Approval): The PR triggers two layers of review. First, an AI review agent runs automatically on the diff, checking Core Craft compliance, security anti-patterns, coding-standard violations, missing test coverage, and architectural drift. This review is advisory: it comments, flags, and may request changes, but cannot approve. FDEs are encouraged to run an agent review on their own work before opening the PR, asking it to critique the change as an adversarial reviewer and catch obvious issues early. The PR reviewer is itself part of the harness, and must be harnessed too. When the FDE or Harness Engineer sees the review agent miss a real issue, raise false positives, or apply the wrong standard, that is a signal to improve the review harness, not a one-off to fix by hand: the correction goes back into the Core Craft review rules (or the review agent’s configuration) so it catches that class of issue automatically next time. Over successive Micro-Cycles the automated review sharpens, shifting more of the burden off humans without lowering the bar. It is the same guide-and-sensor discipline the Harness Engineer applies to execution agents, applied to the reviewer. Second, a human reviews and approves. When team size allows, the reviewer should not be the FDE who built the feature; the natural independent reviewer is the Harness Engineer, whose job is to verify work they did not build. A second pair of eyes catches intent errors, business-logic gaps, and architectural drift that neither pipelines nor AI reviewers reliably detect. Any recurring class of issue the reviewer finds is fed back into the harness (a new Core Craft rule, a new sensor, or a tightened acceptance criterion) so the pipeline catches it next time without a human. In small teams where builder and reviewer are the same person, the same context-switch discipline applies: finish building, then reopen the PR as a reviewer asking “what did the agent get wrong, and what would I challenge if someone else wrote this?” No PR merges without human approval. This is non-negotiable. The AI accelerates the review; the human authorises the merge.
  • SAST (Static Application Security Testing): Automated scan of the source for known vulnerability patterns. If any critical or high-severity finding matches the Blueprint’s threat model, the pipeline fails immediately and the FDE is notified.
  • DAST (Dynamic Application Security Testing): Automated scan of the running application in a sandbox for runtime vulnerabilities (SQL injection, XSS, broken access control).
  • Automated Test Suite: Every acceptance-criteria test the Harness Engineer defined during Blueprinting must pass. No exceptions. A failing test means the feature does not proceed; it goes back to the FDE.
  • Dependency Check: Automated scan for known vulnerabilities in third-party libraries the feature introduces.
  • If Gate 1 passes: The PR merges and the pipeline promotes the build to staging automatically. The cockpit updates to “Staged” and notifies the Product Strategist.
  • If Gate 1 fails: The PR is blocked. The FDE gets a detailed failure report (automated check failures, AI review comments, or human reviewer feedback), fixes the issues, pushes to the same PR branch, and Gate 1 restarts: automated checks rerun and the reviewer is re-notified. No new PR or meeting needed; the cycle is self-service within the existing PR.

LLM-Specific Testing (Agentic Regression Layer): Traditional SAST/DAST and unit tests catch deterministic bugs, not AI-specific regressions. Gate 1 must also include these behavioral checks, especially when the Core Craft or agent prompts have changed since the last deployment:

  • Behavioral regression against golden trajectories: A golden trajectory is a recorded, validated execution trace for a representative task, the sequence of tool calls, decisions, and final output, captured when the agent was known to behave correctly. When the Core Craft or prompts change, the agent’s run is compared against it. Significant deviation (a different tool-call sequence, unexpected file modifications, new external API calls) automatically fails the pipeline and routes to the Harness Engineer. Keep at least 3–5 golden trajectories per major workflow area.
  • Prompt drift detection: Any change to project-instructions.md (the Core Craft) triggers a diff review in the pipeline before merging. It checks for rules that contradict existing ones, removal of security guardrails, and scope changes that expand agent permissions. Prompt drift is the leading cause of production AI regression across enterprise deployments. A change that passes review proceeds to merge; one that fails is escalated as a Flow Sync trigger.
  • Context window pollution check: Verify the Blueprint’s Context Loading Instructions load only what was specified, not the whole repository. If the agent’s actual context token count significantly exceeds the Blueprint’s expected load, it is flagged before execution continues.
  • Adversarial and negative-path testing: Golden trajectories prove the agent succeeds on known-good paths; they do not prove it fails gracefully on unexpected, malformed, or adversarial input. For any Blueprint that processes external data (user input, API responses, ingested logs, support tickets, transcripts), the Harness Engineer sets the adversarial coverage by the Blueprint’s risk profile, the same risk-based approach used for every other verification decision, ranging from a single edge-case check for low-risk Blueprints to a full adversarial suite for Blueprints that ingest uncontrolled external data. Common categories: prompt injection embedded in user-provided data (instructions hidden in a support ticket trying to override behaviour); malformed or schema-violating input that should trigger validation errors rather than silent wrong output; and edge cases where the agent should halt and ask for guidance rather than guess. Keep a small portfolio of failure-mode trajectories alongside the golden ones: recorded traces where the agent is expected to reject input, halt, or surface an error. These negative-path trajectories matter most for Agent-Initiated Blueprint pipelines, where the agent ingests external sources the team does not fully control.

Gate 2: Async Business Verification (Product Strategist).

  • Triggered by (Cockpit): The Product Strategist is notified that the feature is staged and reviews it on their own schedule; no synchronous meeting required.
  • What the Strategist verifies: They review the feature in staging against the Blueprint’s Desired Outcome and Acceptance Scenarios, walking each scenario to confirm it behaves as specified. They are NOT checking code quality; that is Gate 1’s job. They mark the Blueprint “Gate 2 Passed” or “Gate 2 Rejected” on the cockpit with a one-line rationale, within 24 hours of Gate 1 passing to prevent staging queue buildup.
  • User Signal Check: Where relevant, the Strategist may share the staged feature with the FDE’s embedded users or customer contacts for quick informal validation before promoting to production.
  • If Gate 2 passes: The Strategist marks the Blueprint approved on the cockpit and the pipeline promotes the build to production automatically.
  • If Gate 2 fails: The Strategist documents the gap in the Blueprint itself (in planning.md, not a meeting or chat). The Blueprint re-enters the active queue, and the FDE picks it up next Micro-Cycle with a clear written spec of what must change.

4.5 CRAFT Incident Response (Team-Owned, Flow Sync-Driven)

Production incidents are a team responsibility, not an individual one. CRAFT does not introduce a separate incident war-room process; instead, it extends the existing Flow Sync mechanism with an Incident classification and a mandatory post-incident loop. Any team member who detects a production failure, critical regression, or security breach triggers an Incident Flow Sync immediately. No approval is required to call one.

The harness detects incidents too, not just humans. A mature CRAFT system does not wait for a human to notice something is wrong. The Harness Engineer configures monitoring and anomaly-detection agents as part of the observability layer, governance agents whose job is to watch production signals (error-rate spikes, latency regressions, anomalous traces, security alerts, SLO breaches) continuously and around the clock. When such an agent detects a likely incident, it acts automatically: (1) it opens an Incident entry on the State of the Craft cockpit with the detected signal, affected component, and an initial severity estimate; (2) it begins autonomous first-line investigation, correlating logs and traces, identifying the probable failing Blueprint or recent deployment, checking whether a tested rollback path exists in the Wave Manifest, and drafting a preliminary root-cause hypothesis, so that when humans arrive, triage is already underway rather than starting from zero; and (3) it triggers an Incident Flow Sync at the appropriate severity, paging the required humans exactly as a human trigger would. The detection agent investigates and reports; it does not remediate autonomously; containment decisions (rollback vs. hotfix) and any Gate 2 bypass remain documented human calls. This closes the detection gap that pure human monitoring leaves open during off-hours, when the always-on execution model means agents may be shipping and running work while no human is watching. Detection-agent accuracy (true positives vs. false alarms) is reviewed at the CRAFT Evolution and tuned like any other part of the harness.

Incident Classification (set in the cockpit when triggering):

  • P1 (Critical): Service is down or a critical user-facing workflow is fully broken. Trigger the Incident Flow Sync immediately.
  • P2 (High): Core functionality is degraded but the service is operational. Trigger within 30 minutes.
  • P3 (Medium): Non-critical feature impacted. Handle as a fast-tracked Hotfix Blueprint. “Fast-tracked” means it jumps to the top of the next Micro-Cycle’s execution queue ahead of all regular feature Blueprints, but below any active P1 or P2 Hotfix Blueprint. No emergency Flow Sync required, but the Product Strategist must acknowledge the P3 classification on the cockpit before execution begins.

Incident Flow Sync (P1 / P2 only):

  • The triggering team member posts to the cockpit immediately: what is broken, what the impact is, and what has already been tried. They then start the Incident Flow Sync tagging only the required team members.
  • Unlike a standard Flow Sync, the Incident Flow Sync has no fixed time cap; it stays active until the incident is contained. However, every decision must result in an immediate action; discussion without action is not permitted.
  • The Flow Orchestrator leads containment coordination. The FDE who owns the affected Blueprint is the technical authority. The Orchestrator prevents decision paralysis and keeps the team moving with urgency.
  • First decision: rollback or hotfix? If a tested rollback path exists in the Wave Manifest, use it. Rollback is faster and lower-risk. A Hotfix Blueprint is only chosen when rollback would cause greater damage (e.g., rolling back a live schema migration that data already depends on).

Hotfix Blueprint (when rollback is not viable):

  • Follows the same structure as a standard Blueprint but streamlined: replace the Business Goal and Desired Outcome fields with a single Incident Description & Reproduction Steps field. The Harness Engineer’s risk analysis focuses exclusively on the specific failure mode, not a full PRA.
  • Gate 1 still applies, no exception. Under P1 pressure, the automated pipeline runs with a compressed SLA: all acceptance criteria tests must pass. There is no manual override for test failures.
  • Gate 2 is bypassed for P1 incidents. The Flow Orchestrator documents the bypass decision in the Hotfix Blueprint. The Product Strategist reviews the fix asynchronously once the service is restored and confirms it in the cockpit.

Post-Incident Review (blameless, within 24 hours of resolution):

  • 15–20 minutes, facilitated by the Flow Orchestrator. Blameless: the question is never “who caused this?” but always “what in our system allowed this to happen?”
  • Every P1 or P2 incident must produce: a root cause entry in the Local Craft (was this an AI hallucination, a missing HITL checkpoint, an insufficient acceptance criterion, or a process failure?); at least one Core Craft update within 24 hours (the guardrail or rule that would have prevented or caught this); and a follow-up Hardening Blueprint if technical debt contributed to the failure.
  • The incident and its resolution are logged permanently in the State of the Craft cockpit. At the next CRAFT Evolution (Part 1), the team reviews all incidents from the Micro-Cycle to identify systemic patterns, not just isolated fixes.

5. CRAFT Evolution (Sync - Max 50 mins, End of Each Micro-Cycle)

This is the only mandatory synchronous meeting per Micro-Cycle, held at the end of each cycle. (The CRAFT Calibration is also mandatory but runs monthly, see step 6.) It is not a retrospective; it is an active engineering session where the team hardens the system. The Flow Orchestrator facilitates. Attendance is mandatory for the people who build and harden, all FDEs and the Harness Engineer. The Product Strategist joins only for Part 3 (Forecasting & prioritization) and otherwise contributes asynchronously; a shared Strategist or Orchestrator attends the Evolutions of the teams they serve, not one global session.

Agenda (strictly timeboxed):

CRAFT Evolution

6. CRAFT Calibration (Sync - Monthly, Max 2 Hours)

While every other CRAFT event focuses on Technology and Processes, the Calibration is entirely about People. It is held once per month, facilitated by the Flow Orchestrator, and is the only space where team dynamics, workload health, and long-term sustainability are discussed. This event must not be cancelled or shortened; it is the framework’s immune system against burnout and cultural drift.

CRAFT Calibration

The CRAFT Release Strategy: Shipping in Waves

Building a feature is only half the battle. Getting it into the hands of real users, without overwhelming them, breaking existing workflows, or shipping untested code, is where disciplined release management becomes critical. Traditional teams often treat releases as a binary event: the feature is either live or it is not. CRAFT rejects this. In an AI-accelerated environment where features can be built in days instead of weeks, the release pipeline must be just as intentional and structured as the development pipeline. Without a deliberate release strategy, teams risk flooding users with a wall of changes they cannot absorb, burying critical improvements under a pile of minor updates, and eroding the trust they worked so hard to build.

CRAFT manages releases through Release Waves: controlled, phased rollouts that give users time to discover, learn, and adapt to new capabilities before the next wave arrives. Each wave is a curated set of changes, not a dump of everything that passed Gate 2 since the last deployment.

Release Wave Principles

Every release wave is governed by four non-negotiable principles:

  • User Absorption Over Speed: The fact that the team can ship five features in a week does not mean it should. Users need time to discover a new capability, understand how it changes their workflow, and build new habits around it. Shipping too many changes simultaneously guarantees that most of them will be ignored, misunderstood, or resented. A release wave should contain a coherent, digestible set of changes, typically one major feature and a handful of supporting improvements or fixes.
  • Verification Before Velocity: No feature enters a release wave until it has passed the full automated verification pipeline and received explicit human sign-off. AI-generated code is fast, but speed is meaningless if a regression reaches production. Automated tests, security scans, and staging validation are not optional gates; they are the price of admission.
  • Communicate Before You Ship: Users should never be surprised by a change in their workflow. Every wave includes a communication plan: what is changing, why it matters, and where to get help. This applies to internal stakeholders as well: support teams, partner teams, and leadership should know what is coming before it lands.
  • Observe After You Ship: A feature is not “done” when it reaches production. It is done when the team has confirmed, through metrics, logs, and user feedback, that it is working as intended in the real world. Every wave includes a mandatory observation window before the next wave is authorized.

Release Wave Structure

A release wave moves through four distinct phases. The Flow Orchestrator owns the overall wave cadence; the Product Strategist owns the prioritization of what enters each wave.

Phase 1: Wave Staging & Bundling

Not every feature that clears Gate 2 ships immediately. The Product Strategist groups completed features into coherent waves based on user impact, thematic fit, and dependency order. A wave should tell a story: “In this release, users can now do X, and we improved Y to support it.” Random collections of unrelated changes dilute user attention and complicate rollback if something goes wrong.

  • Each wave has a Wave Manifest, a brief document listing every included change, its Blueprint reference, its risk rating, and its rollback plan.
  • High-risk features (new integrations, data model changes, permission changes) are either the sole item in a wave or are gated behind a feature flag for progressive exposure.
  • The Harness Engineer reviews the manifest and confirms that the combined set of changes does not introduce interaction risks that individual Gate 2 checks may have missed.

Phase 2: Automated Verification & Regression Testing

Before any wave reaches a single user, it must pass a comprehensive automated verification suite. This is not a rerun of unit tests; it is a dedicated release verification pipeline that tests the combined behavior of all changes in the wave together.

  • End-to-End Test Suite: Automated browser or API tests that simulate real user journeys across the entire application. AI agents can be instrumental here, so FDEs should instruct their agents to generate and maintain E2E tests as part of every Blueprint execution, not as an afterthought.
  • Regression Test Suite: A curated set of tests that verify existing functionality has not been broken by the new wave. This suite grows with every Micro-Cycle; the Promotion Loop should include rules requiring agents to add regression tests for every bug fix and every critical path.
  • Performance & Load Testing: For waves that include backend changes, automated performance benchmarks catch degradation before users experience it. Thresholds should be defined in the Core Craft so agents bake them in automatically.
  • Security Scan (SAST/DAST): A final security sweep of the combined wave. Even if individual features passed Gate 1 security checks, the combination may expose new attack surfaces.
  • AI-Assisted Test Generation: AI agents excel at generating test cases from Blueprints. Teams should leverage this to maintain high coverage without manual test-writing overhead. The Core Craft should contain rules mandating minimum test coverage thresholds and specifying which test frameworks and patterns agents must use.

Phase 3: Rollout

How a wave reaches users is not a one-size-fits-all decision. A consumer platform serving millions of users has fundamentally different rollout needs than an internal tool used by fifty people, or a startup iterating with a small early-adopter community. CRAFT does not prescribe a single rollout model; it requires the team and organization to deliberately choose a rollout strategy that matches their context, and to document that choice in their Team Agreement so it is applied consistently. (The rollout model is a human release-process decision, not an agent instruction, so it belongs in the Team Agreement, alongside Release Wave cadence, not in the Core Craft.)

The one universal non-negotiable: the team must validate the wave internally before any external user sees it. What happens after that validation is a team or organizational decision. CRAFT provides two reference models; teams may adapt or blend them as needed.

Model A: Direct Rollout (Recommended for: internal tools, small user bases, early-stage products)

After the team has validated the wave in a staging or pre-production environment and confirmed it works in their own workflows, the wave ships directly to all users. This is the fastest path from verification to value, and it is perfectly appropriate when the team has high confidence in their automated test coverage, the user base is small enough that issues surface quickly through direct feedback channels, and the cost of a brief regression is low relative to the cost of delayed delivery.

  • Team Validation (Day 1): The wave deploys to a staging or pre-production environment. The team actively tests the new features in their own workflows. Error rates, performance metrics, and logs are reviewed.
  • General Availability (Day 2+): If team validation passes, the wave rolls out to all users. The communication plan activates: changelogs, in-app announcements, or onboarding nudges go live to help users discover and adopt the new capabilities.

Model B: Tiered Progressive Rollout (Recommended for: large user bases, regulated industries, high-risk changes)

For products with large or diverse user bases, a tiered approach reduces blast radius and provides confidence checkpoints before full exposure. Each tier expands the audience only after the previous tier shows healthy metrics.

  • Tier 1: Internal & Canary (Day 1–2): The wave deploys to the team itself and a small canary group of opted-in users. The team actively uses the new features in their own workflows. Error rates, performance metrics, and user-facing logs are monitored in real time.
  • Tier 2: Early Adopters (Day 3–5): If Tier 1 shows no regressions, the wave expands to a broader early-adopter segment, typically 10–20% of the user base. Support channels are monitored for unexpected friction. The Product Strategist reviews early feedback signals.
  • Tier 3: General Availability (Day 5+): If Tier 2 metrics are healthy, the wave rolls out to all users. The communication plan activates alongside onboarding nudges to help users discover and adopt the new capabilities.

Choosing Your Model: The Product Strategist and Flow Orchestrator should decide on the team's default rollout model during Day 0 setup and record it in the Team Agreement. This default can be overridden per wave: a team that normally uses Model A might escalate to Model B for a particularly high-risk feature (e.g., a payment flow change or a data model migration). The Wave Manifest should explicitly state which model applies to each wave and why.

Emergency Halt (universal, applies to all models): Regardless of the chosen rollout model, the Flow Orchestrator can halt any rollout at any point and trigger an immediate rollback using the pre-defined rollback plan from the Wave Manifest. If error rates spike, performance degrades, or critical user feedback surfaces, the rollout stops. Speed of rollback is as important as speed of delivery. This is not optional; every team, regardless of size or rollout model, must have a tested rollback path before a wave leaves staging.

SLO-Based Automatic Progression (Enterprise Optimization): In enterprise environments, manually gatekeeping each rollout tier creates a bottleneck that negates the speed advantages of agentic development. Teams should define Service Level Objectives (SLOs) for each tier transition: measurable thresholds that, when met, automatically authorize progression to the next tier without requiring human approval. Example SLOs: error rate below 0.5% for 30 consecutive minutes, p95 response latency within 10% of baseline, and zero new critical security alerts. Conversely, if an SLO breach is detected at any tier, the pipeline automatically halts the rollout and triggers an Incident Flow Sync. This SLO-gated automation transforms the tiered rollout from a manually supervised process into a self-governing pipeline: humans define the thresholds, agents monitor and enforce them, and rollbacks execute automatically before degradation reaches end-users at scale. The Flow Orchestrator owns SLO definition and review; SLO thresholds are version-controlled in the Wave Manifest alongside rollback plans. Teams that implement SLO-gated progression consistently report 40–60% faster time-to-general-availability without increasing production incident rates.

Phase 4: Observation Window & Wave Close

After a wave reaches General Availability, the team enters a mandatory observation window (typically 2–3 days) before the next wave is authorized to begin its rollout. During this window:

  • The Product Strategist monitors adoption metrics: Are users actually engaging with the new features? If not, the problem is likely discoverability or communication, not the feature itself.
  • The Harness Engineer monitors error rates and security alerts for any delayed-onset issues that did not surface during progressive rollout.
  • Support and feedback channels are actively triaged. Any bugs discovered are fast-tracked as hotfix Blueprints.
  • Once the observation window closes cleanly, the Flow Orchestrator formally closes the wave on the State of the Craft cockpit and authorizes the next wave to begin staging.

Feature Flags & Progressive Exposure

Feature flags are a first-class citizen in the CRAFT release strategy. They decouple deployment (code reaching production) from release (users seeing the feature). This distinction is critical in an AI-accelerated pipeline where code reaches production far more frequently than users can absorb changes.

  • Every high-risk feature must be wrapped in a feature flag. This is a Core Craft rule, not a suggestion. The FDE's AI agent should be instructed to implement feature flags as part of the Blueprint execution, not bolted on afterward.
  • Flags enable safe experimentation: A/B testing, percentage-based rollouts, and user-segment targeting all become possible when features are flag-gated. This turns every release into a learning opportunity.
  • Flags must have an expiration plan: A feature flag that lives forever becomes technical debt. The Wave Manifest should include a “flag cleanup” date for every flagged feature: the date by which the flag is either permanently enabled and the conditional code removed, or the feature is killed.

AI Agents in the Release Pipeline

AI agents are not limited to writing code; they are equally powerful in the release pipeline when properly instructed. Teams should leverage agents for:

  • Automated Test Generation & Maintenance: Agents can generate E2E tests, regression suites, and edge-case scenarios directly from Blueprint specifications. As features evolve, agents update the corresponding tests, eliminating the manual maintenance burden that causes test suites to rot.
  • Release Note Drafting: Agents can compile Wave Manifests, Blueprint descriptions, and commit messages into user-facing release notes and internal changelogs, saving the Product Strategist hours of manual summarization.
  • Deployment Verification: Post-deployment smoke tests can be generated and executed by agents, providing instant confidence that the deployment was successful before the progressive rollout begins.
  • Anomaly Detection: Agents monitoring production logs and metrics during the observation window can flag anomalies faster than human review, enabling earlier intervention when something goes wrong.
  • Rollback Automation: When a rollback is triggered, agents can execute the pre-defined rollback plan from the Wave Manifest, verify the rollback was clean, and generate an incident report for the team, all within minutes.

The key constraint remains unchanged: humans authorize releases; agents execute and verify them. An AI agent should never autonomously decide to push a wave to General Availability. That decision belongs to the Product Strategist and Flow Orchestrator, informed by the data the agents surface.

CRAFT at Scale: Enterprise Multi-Team Collaboration

In enterprise environments, multiple CRAFT teams often operate within the same complex product. If these lightning-fast teams are not properly aligned, their AI agents will overwrite each other's architectures and create massive technical debt. CRAFT manages scaling through shared context and targeted alignment.

The Shared Core Craft

When multiple teams work on the same product, there is only ONE product backlog and ONE Core Craft file for the entire product repository. All AI agents across all teams must obey this single source of truth. When any single team executes a Promotion Loop during their CRAFT Evolution, that architectural or security update immediately hardens the rules for all other teams working in the repository.

Memory coherence is the multi-team breakpoint. At scale, the single structural failure that most often breaks multi-agent and multi-team setups is not model quality; it is context inconsistency across agents’ memory stores: different agents acting on stale, divergent, or contradictory versions of the same knowledge. The shared Core Craft is CRAFT’s primary defence (one version-controlled source of truth all agents obey), but teams must also ensure that ephemeral, machine-local auto-memories never silently override it (see Personal Rule Files & Auto-Generated Memories) and that promoted rules carry a clear version so every team and agent can confirm they are operating on the current baseline. When agents disagree, the tie-breaker is always the versioned Core Craft, never an individual agent’s private memory.

Cross-Team Alignment Syncs

To prevent collisions and maximize shared learning, humans must align their strategies before the AI generates code, and after it delivers results.

  • Strategist Alignment: Product Strategists from the respective teams collaborate frequently to slice the overarching product vision into independent Blueprints, ensuring no two teams are building overlapping or conflicting features simultaneously.
  • Orchestrator Sync (Platform, Skills & Governance): The Flow Orchestrators across the product group meet regularly to handle both governance and cross-team capability growth. This sync covers: requesting tool licenses from centralized platform teams; standardizing MCP server usage and shared GenAI workflow libraries; sharing new AI Skills, prompt patterns, or automation techniques that individual teams have validated and promoted to their Core Craft; enforcing enterprise data-privacy and compliance regulations; and managing the combined compute footprint and API budget of all AI agents across the product group. Note: In large enterprises with many CRAFT teams, this Orchestrator Sync becomes a critical mechanism for scaling quality. A single team’s breakthrough (a new MCP server, a powerful GenAI workflow, or an effective prompt pattern) can be shared across dozens of teams in a single sync, multiplying its value instantly. Orchestrators should come to this sync prepared with a short list of what their team validated in the last Micro-Cycle that others could benefit from.

Enterprise Promotion: The Open Platform Model

Within a single CRAFT team, the Promotion Loop moves validated knowledge from an individual FDE’s Local Craft into the team’s Core Craft. But in larger organisations, the value of a breakthrough doesn’t stop at the team boundary. A security guardrail discovered by one team is equally valuable to fifty others. A GenAI workflow that halved Blueprint cycle time in one product group should be available to the entire company, not locked inside a single repository.

CRAFT addresses this through a multi-tier promotion model: Local Craft → Team Core Craft → Organisational Craft (or departmental/divisional layers in between, depending on company size). Each tier has its own maintainer and its own review process, but the direction of knowledge flow is always the same: bottom-up.

Why an open platform, not a closed one: Many enterprises today manage development standards through closed, top-down governance bodies: a central architecture board or standards committee that defines rules in isolation and pushes them down to teams as mandates. These closed platforms have a structural flaw: the people writing the rules are disconnected from the people doing the work. Rules arrive late, miss real-world context, and are frequently ignored or worked around because they don’t reflect the reality teams face daily.

CRAFT explicitly rejects this model. The organisational knowledge platform must operate as an open system, following the established InnerSource model, the practice of applying open-source collaboration principles inside the walls of a single organisation. InnerSource gives CRAFT a proven, industry-validated pattern for exactly this problem (used at scale by companies such as PayPal, Bloomberg, and others to let one team contribute directly to another’s shared components, shrinking cross-team timelines from quarters to days). Its core principles map directly onto CRAFT’s multi-tier promotion:

“Open” means open to the organisation, not to the public. This is a critical distinction. CRAFT’s Open Platform Model applies open-source ways of working (anyone can contribute, transparent review, merit-based acceptance, bidirectional flow) strictly within the boundaries of the enterprise. Core Craft rules, security guardrails, GenAI workflows, and the Organisational Craft are shared across internal teams only; they are not published externally, open-sourced to the public, or exposed outside the organisation’s trust boundary. Any decision to release an artifact externally is a separate, deliberate act governed by the organisation’s open-source and IP policies, never a side effect of internal promotion. InnerSource is precisely the term for this: open collaboration, closed perimeter.

  • Anyone can contribute: Any CRAFT team that validates a rule, pattern, or workflow through their Promotion Loop can propose it for promotion to the organisational level. Contributions are not gatekept by hierarchy; they are evaluated on merit and evidence.
  • The Platform Team as Trusted Committer: In InnerSource, contributions from any team are reviewed by a designated Trusted Committer, a maintainer who reviews, mentors contributors, and safeguards quality without inventing the work themselves. In CRAFT, the enterprise Platform Team (or a Platform Owner in smaller organisations) plays exactly this role for the Organisational Craft. They review incoming contributions from every CRAFT team, ensure consistency, resolve conflicts between competing patterns, and maintain the organisational knowledge base, but they do not dictate rules in a vacuum. Critically, the Platform Team reviews at a broader scope than any single team can see: a security guardrail or context pattern that is safe within one team’s architecture may be unsafe when applied across dozens of teams with different tech stacks, data classifications, and compliance obligations. This broader-scope quality review is precisely why the Platform Team layer exists in larger enterprises: it is the safeguard that lets teams move fast locally while keeping the shared platform coherent and safe globally. Reviewing contributions takes real capacity; the Platform Team’s review workload must be explicitly planned for, and expected review timelines must be transparent to contributing teams (an InnerSource discipline, not an afterthought).
  • Bidirectional flow, contribute up and pull hardened rules back down: Knowledge does not just flow one way. A team surfaces a validated rule, the Platform Team reviews and hardens it at organisational scope, and then every team, including the originating one, pulls the reviewed, organisation-grade version back into their environment. What a team gets back is stronger than what it contributed, because it has been stress-tested against the full portfolio. This is the compounding engine of enterprise CRAFT: each team’s local discovery, once hardened by the Platform Team, becomes every team’s baseline.
  • Transparency and traceability: Every rule in the organisational knowledge base has a visible origin: which team proposed it, what evidence supported it, and what problem it solved. Teams adopting organisational rules can trace back to the source and understand why the rule exists, not just that it exists.
  • Opt-in adoption with clear escalation: Organisational rules come in two tiers: mandatory (security-critical, compliance-driven) and recommended (best practices, efficiency patterns). Mandatory rules are enforced via shared pipeline components; recommended rules are available for teams to adopt voluntarily. This prevents the bureaucratic overreach that kills agility in enterprise governance.

The result is a system that hardens from the bottom up: practitioners on the ground discover what works, the Promotion Loop validates it at the team level, the Orchestrator Sync surfaces it across teams, and the Platform Owner curates it into the organisational knowledge base. Every rule in the system was battle-tested before it became a standard, not invented by a committee that hasn’t written production code in years.

CRAFT Enterprise Promotion

Multi-Repository & Monorepo Governance

CRAFT assumes a shared Core Craft for a shared codebase. In enterprise environments, this is rarely that simple. Common scenarios require explicit guidance:

  • Microservices across multiple repositories: Do not put the Core Craft inside a product repository. Instead, create a dedicated governance repository that contains only the Core Craft and shared GenAI Workflows. All product repositories reference it as a Git submodule (or sync via pipeline automation). When a Promotion Loop update is approved, it is merged into the governance repo and automatically propagates to all consuming repos on their next sync. The Orchestrator Sync is the coordination point for this propagation.
  • Monorepos with multiple products or teams: Use a root-level Core Craft for rules that apply universally across all products in the repo, and subdirectory-level Core Craft files for product-specific rules. AI agents load both: the root-level file for universal guardrails and the relevant subdirectory file for product context. This mirrors the cascading context hierarchy already used for Blueprint → Core Craft → Local Craft.
  • Frontend and backend in separate repositories: Treat each as a separate CRAFT team with its own Local Craft and its own contribution to the shared governance Core Craft. Cross-repo Blueprints (features that require simultaneous changes in both repos) must declare their inter-repo dependency in the depends_on field and coordinate execution through the Orchestrator Sync.
  • Rule of thumb: The Core Craft governance structure should mirror your deployment and ownership structure, not your code structure. If two teams deploy independently, they each own a Core Craft. If they share a deployment pipeline, they share a Core Craft.

AI Model Governance

The AI model itself is an operational dependency with a lifecycle, just like any third-party library or infrastructure component. CRAFT teams that do not govern their model usage actively will encounter unexpected regressions when providers deprecate or update models, budget overruns from unoptimized model selection, and audit gaps when they cannot demonstrate which model version produced which output. The Flow Orchestrator owns model governance; the Harness Engineer audits it.

  • Model Version Pinning: Every agent configuration in the team’s tooling must reference a specific, pinned model version, not a floating alias like “latest” or “claude-sonnet.” A floating alias means that a provider-side model update can silently change agent behavior in production overnight with no code change in your repository. Pinned versions must be recorded in the Team Agreement alongside their selection rationale and displayed on the cockpit so the entire team has visibility. When upgrading to a new model version, the upgrade is treated with the same discipline as a Core Craft change: it requires a PR to the Team Agreement, behavioral regression testing against golden trajectories, and the Flow Orchestrator’s explicit approval. Model configuration is an infrastructure decision, not an AI agent instruction, so it does not belong in the Core Craft file itself.
  • Model Selection by Task Type: Not every task requires the most capable or most expensive model. CRAFT teams should define a task-to-model mapping in the Team Agreement. As a reference starting point: use frontier-tier models (the most capable Claude, GPT, and Gemini tier available at the time) for architecture planning, complex multi-file refactoring, threat modeling, and Agentic Forecasting; use mid-tier models for typical feature Blueprint execution and code generation; use lightweight models for test generation, documentation, boilerplate, and release note drafting. This tiered approach can reduce AI compute costs by 40–70% with negligible quality impact if implemented with clear task boundaries. Teams should validate and tune their own mapping; these are starting heuristics, not fixed rules.
  • Model Deprecation Response: When a provider announces a model deprecation, the Flow Orchestrator treats it as a planned maintenance event. The team identifies all pinned references to the deprecated model in their Team Agreement and agent configurations, evaluates available replacement models using golden trajectory comparison, selects a replacement and documents the rationale, and completes the migration before the provider’s end-of-life date. Treat deprecation timelines like security patch deadlines; they are not optional.
  • Model API Availability & Fallback: The Team Agreement should define a fallback model for each primary model in use. If the primary model API returns repeated errors or latency spikes, the agent automatically falls back to the secondary model for the current Blueprint execution. The Flow Orchestrator is notified immediately. Fallback events are logged and reviewed at the next CRAFT Evolution to determine if the primary model is unreliable enough to warrant a permanent switch.
  • Cost Per Blueprint Tracking: Token cost should be tracked at the Blueprint level, not just the team level. When the AI agent executes a Blueprint, the total token spend is logged against that Blueprint’s ID. This enables the Product Strategist to include real AI execution cost in Blueprint ROI calculations and gives the Flow Orchestrator precise data on which Blueprint types are the most expensive to execute, informing context loading optimization and model selection decisions.

Knowledge Threads: Managing AI Context

Generative AI models are highly susceptible to hallucinations when overwhelmed with irrelevant data or starved of necessary context. CRAFT manages this via a strict, cascading hierarchy of markdown files and shared workflows. Agents only read what they need, preventing their context windows from flooding.

Terminology: Skills, Workflows, and MCP Servers: what’s the difference? These three terms appear throughout CRAFT and are sometimes used interchangeably in team conversations. They are distinct:

  • MCP Server (Model Context Protocol Server): A runtime service that exposes tools and data sources to an AI agent via a standardized protocol. Think of it as a plugin your agent can call at execution time: a database query tool, a browser, a file-system interface, a Jira connector. MCP servers give agents real-world capabilities beyond text generation. They are infrastructure components, not instructions.
  • GenAI Workflow: A pre-configured, reusable prompt chain or automation sequence within your AI tooling (e.g., a Cursor workflow, a Windsurf workflow, a custom Copilot agent). A GenAI Workflow is saved as a configured routine that a team member invokes on demand, for example “run the Blueprint review workflow” or “run the release notes draft workflow.” It combines a prompt template with optional tool access and produces a structured output. Workflows live in the team’s shared tooling workspace and are shared across all FDEs.
  • AI Skill (or Prompt Pattern): A validated, reusable instruction pattern or prompt technique that has proven effective for a specific task type, for example a specific way of prompting the agent to write defensive error handling, or a chain-of-thought structure that reduces hallucinations for database schema tasks. Skills are knowledge artifacts, not software components. They live in the Core Craft as rules or in the Local Craft as personal techniques until promoted. When the Promotion Loop promotes a Skill to the Core Craft, it becomes a rule all agents follow automatically, without a human invoking it.
  • Agent-to-Agent Orchestration (Multi-Agent Coordination): As agentic workflows mature, individual agents increasingly need to coordinate with other agents, not just with human operators or external tools. An FDE’s primary coding agent may delegate a subtask to a specialised testing agent, which in turn invokes a security scanning agent, all within a single Blueprint execution. Emerging agent-to-agent protocols (such as A2A and other emerging open standards) enable this by allowing agents to discover each other’s capabilities, negotiate task delegation, and exchange structured results, sitting above the underlying system complexity rather than requiring custom integrations. For CRAFT teams, multi-agent orchestration introduces specific governance requirements: every agent-to-agent delegation must be logged to the same immutable audit trail as human-initiated actions; the Core Craft must define which agents are permitted to invoke which other agents (preventing unconstrained agent proliferation); and the Harness Engineer must include agent-to-agent interaction patterns in their Agentic Chain Verification and governance agent configurations. Multi-agent workflows are powerful force multipliers, but without explicit boundaries they create the same “shadow IT” risk that unchecked SaaS adoption created in the digital era. The Flow Orchestrator tracks agent inventory and inter-agent dependencies as part of their platform governance mandate. CRAFT treats inter-agent trust as zero-trust by default: no agent implicitly trusts another’s output; every delegation and handoff is validated, signed, and written to the same immutable audit trail as human-initiated actions, and the Core Craft explicitly defines which agents may invoke which others. The execution mechanics of coordinating agents (the orchestrator-and-subagent pattern) are covered in the Agentic Execution Discipline section.

The Blueprint (planning.md)

The tactical document for immediate execution. This is the direct output of the Blueprint CRAFTing phase. It is highly explicit, detailing the exact business goal, the desired outcome, what is explicitly out of scope, acceptance scenarios written in plain language, dependency declarations on other Blueprints, the technical How, the Harness Engineer’s risk and verification boundaries (including Product Risk Analysis and threat model), the Agentic Forecast weight, strict Human-in-the-Loop constraints, and a binary, checkable list of acceptance criteria the agent must execute against.

The Core Craft (project-instruction.md / agent.md / Shared Workflows)

The global brain of the product. It sits at the root of the repository and is treated as absolute law by all AI agents. Note: Depending on the Generative AI tooling your team uses, this file may take the name of the tool's default instruction file (e.g., agent.md, claude.md, or .clinerules). If your tool relies on a specific file as its primary source of truth, that file inherently becomes your Core Craft. Most modern AI coding tools support more than just a single instruction file. They provide a rule file system: a directory of individual rule files, each with its own activation mode (e.g., always active, triggered by file-pattern matching, activated by the model’s own judgment, or invoked manually). When splitting the Core Craft into purpose-specific files (as recommended below at the 200-line limit), FDEs must map those files to their tool’s native rule directory structure and configure the correct activation mode for each. A security-rules file, for instance, should be set to always-active; a frontend-patterns file should activate only when frontend code is in context. The Flow Orchestrator is responsible for documenting which rule directory structure and activation modes the team uses, and for reviewing this mapping at every CRAFT Evolution. It contains overarching architecture decisions, strict shared security guardrails (e.g., handling secrets), quality strategies, and mandatory coding standards. This layer also encompasses the shared, centralized GenAI Workflows configured in your tooling for the whole team to use.

The Local Craft (dev-instructions.md & Local Workflows)

The individual Forward Deployed Engineer's sandbox. AI agents operate differently depending on the developer driving them. This file tracks a specific FDE's personal prompting habits, preferred workflow loops, and the unique, recurring mistakes made by their local agent instances. FDEs also experiment with local GenAI tool workflows here before promoting them.

The Promotion Loop (The Mechanism of System Learning)

It is easy to confuse the Promotion Loop with the CRAFT Evolution event, but they are distinctly different: CRAFT Evolution is the meeting, while the Promotion Loop is the specific knowledge management action that takes place during that meeting. It is the mechanical act of hardcoding human learnings into the AI's global brain.

  • Identification: Throughout the cycle, if a Forward Deployed Engineer notices their AI consistently hallucinates a specific API call, attempts to bypass security, or executes highly inefficient loops, they document a corrective rule in their personal Local Craft. Similarly, they may build an effective new local GenAI workflow. The Promotion Loop is not limited to AI rule fixes; it also applies to human learnings: a clever architectural approach, a problem-solving technique, a user-interaction insight, or a workflow pattern that proved significantly more effective than the team’s current default. These are equally promotable.
  • Validation & Inspection: During the synchronous CRAFT Evolution event, developers share these local insights. The team evaluates if the AI mistake is a systemic risk or if the new workflow is valuable enough for the whole team to adopt.
  • Promotion & Cleanup: If validated, the constraint (or tool workflow) is physically copied out of the Local Craft and pasted into the Core Craft (or configured in the shared tool workspace). Crucially, the rule or workflow must then be deleted from the local setup to maintain a single source of truth and prevent conflicting instructions. This cleanup extends to the tool’s own rule file hierarchy: if the promoted rule existed as a personal or workspace-scoped rule file in the FDE’s local tool configuration, that file must also be removed or updated to reference the Core Craft version. Similarly, if the insight originated from the tool’s auto-generated memory store, the FDE should verify that the memory does not conflict with the newly promoted Core Craft rule; stale or contradictory auto-memories are a common source of agent confusion after promotions.
  • Systemic Hardening: Because all agents are instructed to treat the Core Craft as absolute law, that specific mistake is instantly eradicated across the entire team, and new efficiencies are immediately scaled. The framework learns, adapts, and hardens in real-time.

CRAFT Promotion Loop

The Team Agreement

The Team Agreement is the team’s operational contract: a living document that captures all process decisions, governance policies, and infrastructure configurations that govern how the humans work together, but that AI agents do not need to read. It is explicitly not part of the Core Craft and is never loaded into an agent’s context window. Where the Core Craft tells the AI how to behave, the Team Agreement tells the humans how to operate.

What belongs in the Team Agreement:

  • Micro-Cycle duration (e.g., 5 days) and the rationale behind it.
  • Model governance: pinned model versions, task-to-model mappings, fallback model configurations, and model deprecation response plans.
  • Budget and cost targets: API token budgets per Micro-Cycle, cost-per-Blueprint thresholds, and escalation rules when budgets are exceeded.
  • Meeting cadences and timeboxes: any team-specific adjustments to Evolution, Calibration, or Flow Sync defaults.
  • Release Wave cadence & rollout model: how frequently the team ships waves, the default rollout strategy (e.g., direct vs. tiered/canary), and any user-absorption constraints.
  • Cockpit configuration: which tool the team uses, what custom fields are tracked, and how automated state transitions are wired.
  • Team composition and role assignments: who holds which CRAFT role, and any role-combining decisions for small teams.

Visibility: Key operational parameters from the Team Agreement, especially Micro-Cycle duration, current model versions, and budget status, should be surfaced on the State of the Craft cockpit so they are visible to the entire team in real time without opening a separate document.

Ownership: The Flow Orchestrator maintains the Team Agreement and reviews it at every CRAFT Calibration. Changes follow the same PR discipline as the Core Craft: proposed via pull request, reviewed by at least one other team member, and merged with a clear rationale. The Team Agreement is version-controlled alongside the codebase but is never referenced by any AI agent configuration.

The distinction matters: The Core Craft has a strict 200-line budget because every line competes for the AI agent’s attention. Filling it with operational governance (model versions, budget limits, meeting schedules) wastes that budget on information the agent will never act on, while pushing out the security guardrails and coding standards it must follow. Keep the Core Craft for agent instructions. Keep the Team Agreement for human agreements.

Measuring CRAFT Success

CRAFT does not use velocity points or sprint burndowns. Success is measured by the health, speed, and sustainability of the overall system. If these signals are trending in the right direction, the team is executing CRAFT correctly.

The Meta-Signal: Agentic Forecasting Accuracy. Of all the metrics below, one stands above the rest as a composite indicator of overall team maturity: the accuracy of the AI-generated Effort Weights from Agentic Forecasting. When forecasting accuracy improves over time, it means the Core Craft context is rich and current, Blueprint specifications are precise enough for the AI to reason about scope, the team’s execution patterns are stable and predictable, and the codebase is well-structured enough for the agent to estimate impact reliably. Conversely, declining forecasting accuracy is an early warning that one or more of these foundations is degrading, before the impact shows up in cycle time or quality metrics. The Product Strategist and Flow Orchestrator should treat Effort Weight Accuracy (tracked under Flow Metrics) as the single best leading indicator of system health: when the AI can accurately predict how long work will take, the entire CRAFT system is functioning well.

Implementing CRAFT: Getting Started

Adopting CRAFT is not a big-bang transformation. The framework is designed to be bootstrapped incrementally, starting with the foundational infrastructure and letting the process harden organically through each Micro-Cycle. Below is the recommended sequence for a team adopting CRAFT for the first time.

Day 0: Build the Foundation Before You Build Anything Else

Before a single line of AI-generated code is written, the team must establish the three non-negotiable foundations:

  • Stand up the Core Craft: The Flow Orchestrator and FDEs collaborate to create the initial project-instructions.md (or tool-equivalent) at the root of the repository. At a minimum, it must define: the tech stack and architecture principles, security non-negotiables (secrets handling, forbidden patterns), coding style standards, and which HITL checkpoints apply globally. This does not need to be perfect; it needs to exist. It will harden with every Micro-Cycle. Rather than starting from a blank file, teams can bootstrap the Core Craft from a persona-based agent framework (e.g., BMAD or RuFlo; see “Bootstrapping with Persona-Based Agent Frameworks” below). Installing such a framework first gives you a working set of agent personas, workflow templates, and default rules; the team then reviews, trims, and maps those defaults into the Core Craft, so the Core Craft starts already grounded in a proven baseline rather than being written from scratch.
  • Set up the State of the Craft cockpit: The Flow Orchestrator configures the team's live tracking tool. Blueprint states, pipeline health signals, and API budget visibility must be in place before work begins. A team flying blind on Day 1 will never build the habit of trusting the cockpit.
  • Establish the AI toolchain and MCP servers: The FDEs confirm which Generative AI tools, MCP servers, and GenAI workflows are available. Any missing access (licenses, credentials, enterprise approvals) is flagged immediately to the Flow Orchestrator to unblock before development starts, not after.

Brownfield Adoption: Starting on an Existing Codebase

Day 0 assumes a new project. Most enterprise teams are not starting from zero; they have a 3-year-old Rails monolith, a microservices ecosystem with undocumented conventions, and a team already running Scrum. CRAFT is explicitly designed to be adopted incrementally into existing environments. Do not attempt a big-bang transformation.

Stage 1: Build the Core Craft from what exists (Week 1–2)

  • Use an AI agent to analyse the existing codebase and generate a first-draft Core Craft. Feed it the repository, any existing architecture docs, coding guidelines, and security policies. The output will be imperfect, so the team reviews and trims it to the 200-line maximum, removing anything stale or incorrect. This is faster than writing from scratch and more accurate than doing it from memory.
  • Identify the existing test suite. If automated coverage is below 60%, designate the first Hardening Blueprints of the adoption to close that gap before AI-generated features rely on it.
  • Identify the most critical HITL thresholds for this codebase (e.g., no AI modifications to payment logic, no schema migrations without Harness Engineer sign-off). Add these to the Core Craft first; they are non-negotiable before any AI agent touches the codebase.

Stage 2: Run CRAFT on new features only (First 2–4 Micro-Cycles)

  • Do not convert the existing backlog to Blueprints overnight. Pick the next 2–3 new features and run the full Blueprint CRAFTing process for those. Let the existing team handle existing bugs and maintenance in their current way while the team learns the CRAFT rhythm.
  • Run the CRAFT Evolution after each Micro-Cycle even if the cycle was short or imperfect. The Evolution is where the team learns; skipping it defeats the purpose.

Stage 3: Full migration (Micro-Cycle 5+)

  • Once the team has run 3–4 full Micro-Cycles with confidence, migrate all new work to Blueprints. Legacy bug fixes and maintenance can continue informally for now, but new features and significant enhancements must follow the full CRAFT flow.
  • A realistic timeline to stable CRAFT operation on a brownfield codebase: 4–6 Micro-Cycles. Teams that try to rush this consistently report higher Core Craft quality issues and lower Gate 1 pass rates in the first few months. Patience in the adoption stage compounds into velocity later.

Agent-Driven Team Onboarding

When a new team member joins a CRAFT team, they should be onboarded by an agent, not by scheduling a series of knowledge-transfer meetings. The team maintains an onboarding.md workflow file, an agent-executable document that a new team member runs with their AI agent on their first day.

The onboarding.md workflow typically instructs the agent to: walk the new member through the full Core Craft (reading each rule and explaining the reasoning behind it), review the team’s current HITL thresholds and what triggered their creation, summarise the Blueprints currently in progress and their status on the cockpit, explain the Local Craft conventions each FDE follows, and run a dry-run Blueprint review exercise using a recently completed Blueprint as a worked example. The agent asks questions and checks understanding rather than just presenting information passively.

Maintaining the onboarding workflow is the Flow Orchestrator’s responsibility. The onboarding.md is reviewed at every CRAFT Calibration (Part 1: Framework Adoption Health) and updated whenever: the Core Craft has a major structural change, a new tool or MCP server is added to the team’s toolchain, HITL thresholds change, or the team receives feedback that a new member was confused about something the workflow should have covered. An outdated onboarding workflow is an impediment; the Orchestrator treats it like any other blocker in the system.

Skills for CRAFT Teams

A common question when adopting CRAFT is: what skills do our people actually need? The honest answer is that the scarce skill has moved upstream. Value no longer comes from typing code (agents do that cheaply) but from directing and verifying the system that produces it. Every CRAFT team member makes the transition from code generator to system steward and verifier. The skills below reflect that shift. They are deliberately stack-agnostic: CRAFT prescribes no language, framework, or model, so every skill is framed as “strong fundamentals in your team’s chosen tools,” never fluency in a specific one.

First Micro-Cycle: Learn by Doing, Not by Planning

The best way to learn CRAFT is to run a real Micro-Cycle on a scoped, low-risk feature. Deliberately choose something that touches all four roles so every team member experiences the full loop end-to-end.

  • Forward Deployed Engineer: Draft the first real Blueprint in planning.md from a genuine user need, then add the technical layer, configure their Local Craft, set HITL thresholds in the AI tooling, and execute the Blueprint with the AI agent. Resist the urge to over-specify; a good first Blueprint is two pages maximum. Log every hallucination, every HITL halt, every prompt that worked or failed.
  • Product Strategist: Refine the Blueprint’s business Why, confirm the desired outcome and scope, prioritize it, and later run Gate 2 business verification. Practice owning the decision of what to build and when, not authoring the technical detail.
  • Harness Engineer: Conduct the first PRA and add the initial acceptance criteria. The CI/CD pipeline gates do not need to be complete on Day 1, but at least one automated test must run as proof of concept.
  • Flow Orchestrator: Observe the team's first execution actively. Note every friction point. At the end of the cycle, facilitate the first CRAFT Evolution with the specific goal of producing at least three Core Craft rules from the team's logs.

After the first Micro-Cycle, the team should hold a short informal debrief (separate from the formal Evolution) to answer: What felt right? What felt like we were fighting the framework? What would we do differently? These answers refine the next cycle.

CRAFT Anti-Patterns: What to Avoid

Most teams that struggle with CRAFT are not failing because of the technology; they are failing because they unconsciously drag old Agile habits into the new system. These are the most common anti-patterns and how to recognize them.

Agentic Execution Discipline: Practical Guidance for FDEs

The following principles are not part of the CRAFT Engine’s workflow phases; they are practical execution habits that every Forward Deployed Engineer must internalise. Mastering these disciplines is what separates an FDE who uses AI from an FDE who orchestrates AI effectively.

The Next Step: Self-Hardening Agents

Everything above places a CRAFT team at what the wider field calls Orchestrated: agents execute and coordinate the work, and humans govern the system and harden the harness by hand. The Harness Engineer runs the root-cause analysis, decides what becomes a Core Craft rule, and keeps the Promotion Loop turning. The learning is real, but it is human-run. There is one more step, and it does not replace anything here: it adds a class of agents that take on part of the hardening itself, under supervision.

The principle is the one CRAFT already lives by, System Over Output. When something breaks in production, the reflex is to fix the output; a CRAFT team fixes the system so the whole class of failure cannot recur. Self-hardening agents automate the first half of that loop. A Hardening Agent watches the signals the team already produces, bug reports, incident alerts, error traces, failing evals, and when it finds a recurring or high-confidence issue it does two things in the background: it diagnoses and drafts a fix for the output, and it drafts the matching harness change, a new eval, a Core Craft rule, or a guardrail, so the same mistake is caught next time. It arrives not with a patch, but with a patch and a reason it will not happen again.

What keeps this safe is the discipline that already defines CRAFT: the human stays in the loop. The agent proposes; an engineer reviews and accepts or rejects. Nothing it drafts, neither the output fix nor the harness change, merges on its own. Each change clears the same Gate 1 and the same human review as any Blueprint, and the agent that drafts a change is never the instance that approves it (the propose-and-verify separation from the roles section holds here too). What the agent removes is the toil of noticing, diagnosing and drafting; what the human keeps is the judgement and the final say.

This is the move from Orchestrated toward AI-native: the system begins to improve itself rather than waiting for a person to improve it, which is the self-learning loop that defines the top of the maturity ladder. But it is AI-native with the accountability left in. Because an engineer accepts every change, a named human still answers for what ships and for how the system evolves. Cross that line, let the hardening agents merge their own changes with no human approval, and you reach the fully self-governing corner CRAFT deliberately does not recommend today, for the same reason it keeps humans on every other gate: an agent cannot be accountable, and hardening the system is one of the highest-stakes things a team does. Where a given CRAFT team lands, Orchestrated or the supervised edge of AI-native, is therefore an implementation choice. It depends on how much of the hardening loop you let the agents run, and you should widen that only as fast as your evals, observability and review capacity can keep the changes safe.

The CRAFT Commitment

CRAFT is not a process to be adopted partially. It is a system designed around one principle: that the combination of a small, expert human team and a well-instructed AI engine will consistently outperform any traditional development organization of any size. But this only holds true when the system is operated with discipline.

The Core Craft must be maintained. Blueprints must be explicit. HITL thresholds must be respected. The Promotion Loop must run. The Calibration must not be skipped. Anti-patterns must be called out and corrected by the team, not tolerated for the sake of short-term speed.

Teams that commit to this discipline will find that, within a few Micro-Cycles, the system begins to compound. Each Evolution hardens the AI's behavior. Each Orchestrator Sync spreads that improvement across the organization. Each Calibration strengthens the team's ability to sustain the pace. The result is not just faster software; it is a continuously self-improving engine for delivering business value, built on a foundation of human expertise, AI execution, and shared accountability.

The expert is always in the lead. The AI accelerates the execution. And the team’s real product is the system that produces the output, not the output itself. CRAFT makes all three possible at scale.

Version: 1.1
Released: April 2nd, 2026 (last updated July 2026).
Feedback? Please use the contact form!

Key Terms: CRAFT Glossary

A quick reference for the vocabulary used throughout this framework. Expand to scan the terms an implementer needs to hold in mind.

Blueprint
The single source of truth for one unit of work (planning.md): a machine-readable, AI-executable spec with business goal, desired outcome, scope, acceptance criteria, technical How, risk boundaries, HITL thresholds, and Effort Weight.
Core Craft
The global “brain” of the product (e.g. project-instructions.md / agent.md): the shared rules, guardrails, and standards every AI agent treats as absolute law. Hard-capped at 200 lines.
Local Craft
An individual FDE’s personal instruction file (dev-instructions.md): their prompting habits and the recurring mistakes of their local agent, before promotion.
Team / Organisational Craft
Higher promotion tiers: validated rules flow Local → Team Core Craft → Organisational Craft via the Promotion Loop and the Platform Team (InnerSource).
Harness
Everything wrapping the AI agent: instructions, context, tools, guardrails, tests, and observability. Split into guides (steer before it acts) and sensors (verify after it acts).
Micro-Cycle
The team-level cadence period (typically 3–7 days) whose boundary triggers the CRAFT Evolution. Not a per-Blueprint sprint.
Gate 1 / Gate 2
The two verification gates: Gate 1 = automated technical verification + independent human PR review (Harness Engineer); Gate 2 = async business verification (Product Strategist).
Effort Weight
A normalized AI-generated complexity score (not story points) produced by Agentic Forecasting; combined with business value to rank the backlog.
HITL
Human-in-the-Loop: explicit stopping points where the agent must pause for human approval, the technical expression of “Expert in the Lead.”
Flow Sync
A max-15-minute ad-hoc intervention to unblock the system, not a standup. Has an Incident variant for production failures.
CRAFT Evolution
The one mandatory sync per Micro-Cycle (max 50 min) where the team hardens the system and runs the Promotion Loop.
CRAFT Calibration
The monthly people-focused sync (max 2h): adoption health, workload, wellbeing, and sustainability.
Promotion Loop
The mechanism that moves a validated learning out of the Local Craft into the Core Craft (and up the tiers), deleting the local copy to keep one source of truth.
Release Wave
A curated, phased rollout of completed features to users (deployment batching). Distinct from Wave-Based Execution (decomposing one Blueprint’s execution).
Golden Trajectory
A recorded, validated agent execution trace used as a behavioural-regression baseline in Gate 1.
State of the Craft (Cockpit)
The always-live operational view that replaces the standup: Blueprint status, agent activity, pipeline health, and cost.
Team Agreement
The humans’ operational contract (cadence, model governance, budgets), deliberately NOT part of the Core Craft and never read by agents.

CRAFT FAQ

Common questions from teams evaluating or adopting the CRAFT framework.

The three phases

CRAFT

Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.

ING

Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.

AI

Continuously assess and refine the output based on the prompts output to improve the overall quality.