v1.4.0next
Get started

Why Kraft

What Kraft does that an interactive session, a loop, or a skill does not, and when it is the wrong tool.

Kraft is not a coding agent, and it doesn't replace yours. It runs the agent you already use (Claude Code, Codex and others) and adds what a single agent session can't give you.

Use Kraft when you want agent work to run while you are not watching, and to stop only where a person should decide. Use an interactive session for everything else. This page explains what Kraft adds, what it does not, and how it compares with the tools you already use.

What your agent can't do on its own

  • A process written once. A chain is a file you can read and diff: spec, plan, implementation, verification, review, merge request, merge. The agent does not decide the workflow fresh each time.
  • Gates where a human decides. You approve a spec, a plan, or a merge request. A rejection carries your note back to the node that wrote it, and that node tries again. See Gate.
  • Checks the agent doesn't grade itself on. An agent that says "all tests pass" is reporting on its own work. Kraft runs your tests in a node of its own, waits for your CI itself, and has a separate reviewer or judge decide whether a fix is done. A worker can't approve its own gate.
  • Bounded retries and spend. Time, token, and dollar caps apply per work item, node, step, and task. A defect the agent cannot fix stops the item and shows you the full trace. See Caps and budgets.
  • Answers to permission prompts. A worker runs with nobody to click "allow". Kraft answers from a policy you set and records each allow and deny on the item's timeline. See Why a permission gate.
  • Isolation. Every work item edits its own git worktree on its own branch, so your checkout stays as you left it and items run in parallel.
  • One board. All your repos in one view, grouped by what needs you, what is running, and what is done. Search and analytics (lead time, cost by node) come from the same events. You can reach the board from a phone through a tunnel; see Remote access.
  • Several agents and models in one work item. Claude Code, Codex, Cursor, Gemini, OpenCode, and Amp run as harnesses, and each agent task picks its own harness and model. One item can have one agent write the code and another review it, or a strong model plan and a fast one triage. The shipped chains run every agent task on Claude Code; to mix them, change the tasks' harness: in library.yaml. See Switch a task to another harness.
  • Several ways in. File work from the UI, the CLI, an agent over MCP, or a trigger on a schedule or a webhook. A work item lands paused and spends nothing until a person starts it. You can also start it as you file it: the UI's Create and start, or kraft item create --autostart. Only auto-intake, which is off by default, starts work on its own.

How it compares

An interactive session

In a session you are the scheduler, the reviewer, the test runner, and the memory. Close the terminal and the work stops. Kraft takes over those jobs: items run unattended, several at once, and the board tells you which ones need you.

A loop

A loop (a while loop around an agent, a Ralph-style run, /loop) repeats until the agent decides it is done, and nothing bounds that decision. It spends tokens on a defect it cannot fix and has no way to ask a person. Kraft's retry loops stop at an attempt limit, a wall-clock limit, or a token or dollar budget, and then escalate to you.

Skills and plugins alone

A skill tells an agent how to do one step well, and the agent still decides whether to run it, in what order, and when to stop. That has two costs:

  • It is inconsistent. The process lives in the agent's judgment, so the same request can skip a review one day and run it the next. Nothing stops at your approval unless the agent chooses to, and nothing resumes after a crashed session. In a chain the order, the gates, and the retries are fixed by the file, not decided at run time.
  • It is expensive. Every step, tool call, and result stays in one growing context that is paid for again on each turn. In Kraft the chain, not the agent, does the sequencing and bookkeeping. Kraft runs the repo's tests, the rebase, and the merge-request steps itself (builtin and forge tasks, see Task), and each agent task is its own headless session, not another turn in one long conversation.

Kraft runs your skills as tasks inside a chain. Kraft Lite runs the same chain inside one agent session when you do not want a service. It keeps the fixed order and gates, but not the fresh-session saving, because it runs in the one session.

CI

CI runs a fixed script on a fresh runner. Kraft runs an agent on your machine that you can pause, steer, or reject with a note. Kraft does not replace CI. It watches your merge request's pipeline and fixes what fails.

When not to use Kraft

  • A quick fix. Open a session. A chain is overhead you won't earn back.
  • Exploration. If you don't yet know what you want, you need a conversation, not a pipeline. Explore first, then hand off a spec.
  • Work you want to steer every few minutes. That is what a session is for.
  • A shared team service. Kraft runs on one machine and binds loopback by default. It is not a hosted service.
  • A repo that isn't set up for it. Kraft stops a repo's first item rather than guess its test_command or setup_command. See Troubleshooting.

The usual pattern uses both: think in a session, settle a spec and a plan, hand them to Kraft, and take the next thing.

Next steps

Copyright © 2026