nextv1.4.0
Concepts

Caps and budgets

How Kraft bounds time, tokens, and dollars, and how a policy resolves from policy.yaml down to a single task.

Kraft stops a work item that runs too long or spends too much, and it hands the item to you instead of failing it. This page explains how those bounds fit together and why they work the way they do. For every key and its default, see Configuration.

Two kinds of bound

Kraft bounds work in two ways.

Loops bound repetition. A fix loop retries a node until it passes. Its attempt limit is max_attempts in the chain (attempts under loops: and default: in policy.yaml), and its wall clock is wall_clock_s there. An external wait (a pipeline, an automated review, an approval, a merge landing) is not a loop. Its timeout is its own task's total_time_cap_minutes.

Caps bound a scope's resources: time, tokens and dollars. A scope is the whole work item, one node, one step or one task.

Only a person resets a loop's count. A retry from the board or from kraft item retry starts the retried span's fix loops afresh. A retry that a worker, an MCP assistant, an escalation turn or a rate-limit relaunch asks for keeps counting. An escalation turn is the agent Kraft starts to help when an item stops for a person; see Fallback and escalation.

Why loops: ships empty

A loops: key that names no live loop does nothing and warns nobody. A stale cap would look live and bind nothing, so the shipped policy.yaml names none. The default: entry covers every fix loop. A loop's own max_attempts in the chain wins over default:, so the shipped fix loops take their attempts from the chain and only their wall clock from default:.

Caps are set per level

You set the four caps (time_cap_minutes, total_time_cap_minutes, token_budget, budget_usd) per level of scope: the work item, its nodes (a gate is a node), its steps and its tasks.

A level's default binds every scope of that kind that nothing set a value for. The shipped policy.yaml has example defaults, commented out. If you uncomment them, the work item runs under 480 minutes, each node under 180, each step under 120 and each task under 90, all at once. Out of the box no level default is set: only the caps the shipped tasks set on themselves in library.yaml apply, alongside the budget: dollar caps. Whichever runs out first stops the item, and the stop names its scope.

A value set in the chain (on the chain, a node, a step or a task) or by the item's own override replaces the defaults for that scope and for every scope under it that sets none. A node set to 240 minutes runs its steps and tasks under 240 too.

A default is not a ceiling. Any scope may set more, up to its level's maximum. A narrower level's cap may not exceed a broader one's (tasks <= steps <= nodes <= work_item), so a task cap of 20 minutes inside a step capped at 10 is refused when the chain loads.

A repository's caps (its policy: in repos.yaml) are the work item's, like the chain's own policy:.

Time caps

There are two clocks.

  • time_cap_minutes counts running time. Paused time, external waits, gates and rate-limited time do not count.
  • total_time_cap_minutes counts wall-clock time, including waits, gates and rate limits, less only a manual pause. For a task other than a wait, it is its running time. For a wait, it is the wait's timeout.

Every clock counts from the scope's start in the current run. A person's retry starts them all afresh. A rate-limit relaunch or a stuck escalation's own retry does not.

A node's clock is its own, independent of its fix loop's timeout_minutes, which bounds only the fix cycles. A node's cap covers everything that runs as that node: its tasks, recovery, fix loop, a gate's automated review and the automatic stuck-escalation turn. Your own escalation chat is not the node running, so it is not capped.

Hitting a cap stops the item for you with the reason "<scope> hit its time cap of N minutes". It is never treated as a code failure. It spends no recovery or fix-loop attempt, and no stuck escalation answers it. A gate's own timeout stops the item the same way, naming the gate, and the gate stays open for you to approve or reject.

Because a work item's total_time_cap_minutes counts its gates and external waits, a cap of two days stops an item still waiting on a week-long approval. Size it for the longest wait your chain allows.

Token and dollar budgets

A token or dollar cap refuses to start the next agent launch once its scope, or any scope around it, has reached its cap. Kraft checks before each launch, gate reviewers and escalation turns included. It cannot interrupt a running agent, so overshoot is bounded by one launch's cost. While that agent is still running, its spend counts too: Kraft prices its known tokens against a packaged rate table and shows the result as an estimate (~$4.20 (est.)) until the session exits and the agent's own reported figure replaces it — so a long session's spend can trip a dollar cap before it finishes, not only after.

Three things follow.

  • A finished launch whose harness reported no cost counts at an estimate from the same rate table. If Kraft cannot price it, the caps differ: see Harnesses that report no cost.
  • Launches that start together in one scope, such as a parallel step's tasks, each pass the check before any of them has spent. A shared scope's cap can be overshot by up to their combined cost.
  • A launch that resumes an earlier agent session counts only what it added.

budget.work_item_usd and budget.daily_usd in policy.yaml stay the outer ceilings, above any per-scope budget_usd.

Harnesses that report no cost

Codex, Cursor, Amp and Antigravity report tokens but no cost. Kraft estimates a finished session's cost from prices.json, the rate table packaged with each release (Anthropic, OpenAI and Google models, refreshed per release), priced on the model the log names or, when it names none, the model Kraft launched it with:

  • Codex names no model in its log, so it is priced on its launch model (gpt-5.6-sol in the shipped deep profile). A task that passes no model runs on the CLI's own default, which Kraft cannot price.
  • Amp names its model in its log, and is priced when the table lists it.
  • Cursor reports its model as "Auto", so it is priced only on a launch model the table lists. Cursor's own model ids usually aren't. A listed one, such as gpt-5.6-sol, is priced at its provider's list price (OpenAI's, here), not what Cursor bills you for it on your plan or in credits.
  • Antigravity names its model in its log only when launched with one. It is priced when that is a base model id the table lists, such as gemini-3.8-flash with effort set, never a slug with the effort in it (gemini-3.8-flash-low). That is Google's API list price, not the quota or credits a signed-in account spends.

A session Kraft cannot price has no dollar figure, and the two kinds of dollar cap treat it differently:

  • budget.work_item_usd, budget.daily_usd and kraft item create --budget count it as $0 and never stop for it. The item's timeline says so the first time (spend_unpriced), and kraft admin doctor warns about each harness the chains launch on an unpriced model.
  • A per-scope budget_usd counts it as unknown spend and stops the next launch in that scope for you.

The estimate stays marked as one (~$4.20 (est.)), never as the agent's own figure. Some models cost more once a request's context passes a size, such as OpenAI's past 272k tokens. Kraft prices a whole session at that higher rate when one of its requests passed the size, which it can tell only from a log that shows each request: Claude's and Amp's. Codex's log shows only running totals, so a Codex session is priced at the base rate, and a long-context one estimates low.

A dollar cap does not bound a session Kraft cannot price. token_budget bounds any harness.
You want to boundClaude, OpenCode (report cost)Codex, Amp, Antigravity on a priced model (estimated)Cursor, or any unpriced model
One work itembudget.work_item_usd, --budget, or budget_usd on the work itemThe sametoken_budget on the work item
One node, step or taskbudget_usd or token_budget on that scopeThe sametoken_budget on that scope
Everything in a daybudget.daily_usdbudget.daily_usdNo cap counts it.

Today's spend against that cap is GET /budget/today, and every item's own share of it is budget_cap.daily on GET /work-items/{id} (see HTTP API).

Raising a cap on a stuck item

A work item's own cap can exceed the chain's, up to maxima.work_item. To unstick a capped item, raise its cap and retry. You need no config edit:

kraft item set-policy ID --policy time_cap_minutes=240
kraft item retry ID

set-policy replaces the item's whole override, so repeat every --policy you set earlier, or they are dropped. --clear removes the override. A cap on one path, such as verification.time_cap_minutes=90, may exceed that scope's level default up to its level's maximum, but not what the chain, the item's cap or an enclosing path already gives it. Kraft refuses a larger value and names both scopes.

An item has two dollar caps, and each has its own way up:

  • budget_usd in the item's policy is a per-scope cap, like the time caps above. set-policy sets it the same way: item-wide it may exceed the chain's up to maxima.work_item, and on a path it only tightens. Then kraft item retry.
  • The item's own dollar cap, the one kraft item create --budget sets in place of budget.work_item_usd, is raised with kraft item raise-budget ID --usd N (the MCP raise_budget tool, or the Raise budget button on the item), which also retries the item. --usd none lifts the cap. It works only on an item this cap stopped, and no maxima bounds it. An item stopped by budget_usd, token_budget or budget.daily_usd shows no Raise budget, and raise-budget refuses it: raise budget_usd or token_budget with set-policy as above or in policy.yaml, and budget.daily_usd in policy.yaml, or wait for local midnight. Then kraft item retry.
kraft item raise-budget ID --usd 25

A budget_usd stop on unknown spend is different: no higher cap passes it, since Kraft still cannot show the scope is under the new one. Clear the item-wide cap instead, then retry:

kraft item set-policy ID --policy budget_usd=none
kraft item retry ID

budget_usd=none means no dollar cap for the work item, and none for each node, step or task the chain set no budget_usd on, replacing the chain's, the repository's and the policy.yaml defaults. It is item-wide only, and refused when maxima.work_item.budget_usd is set. A cap the chain set on a node, step or task still binds; skip that node with kraft item skip instead. token_budget, budget.work_item_usd and budget.daily_usd still bound the item.

The shipped implementer cap

The implementer task, which both shipped chains run, sets time_cap_minutes: 120. In practice, successful implementer runs finish well inside that time, so the cap binds only a run that has already gone wrong, for example a worker stuck waiting on a test run it backgrounded.

A long plan can need more. Raise the item's own cap, or set a larger one on the task in library.yaml. The task's 120 exceeds the example's tasks default of 90, which a scope may do. A maxima.tasks.time_cap_minutes (or a broader level's) under 120 refuses both shipped chains until you lower the task's cap to fit.

Background jobs

A worker's session ends with its turn. If the turn ends with a background job still running, the session fails at once, and the log and a background_jobs_abandoned event name each job.

How a policy resolves

A work item freezes its policy when you file it. Kraft layers it from the broadest source to the narrowest:

  1. defaults and maxima in policy.yaml.
  2. The repository's policy: in repos.yaml.
  3. The chain's policy:.
  4. Each node's, step's and task's own policy:.
  5. Your item's own override, from kraft item create --policy or kraft item set-policy.

Kraft refuses at intake a layer that relaxes what it inherits, and names the scope. That is why a bad policy fails when you file the item, not hours into a run.

Your override comes last, so its operational values win over the template's. A max_attempts or timeout_minutes on an execution node wins over its fix loop's own max_attempts. Its safety values only ever tighten. An allowed_tools list intersects with what the scope already allows, a grants list keeps only the grants the scope already holds, deny_tools adds to the list, and a sandbox must match any already set. The chain narrowing the same field first never causes a refusal.

Your override is the one layer that can change after filing. On a running or waiting item, a change takes effect from the next node entered and the next check of a wait.

A recovery plan inherits the task, step or node that declares it. A fix loop, its judge, a node's escalation task and its conflict handler inherit their node. A gate's auto_review inherits its gate.

Why maxima start unset

A fresh install ships with defaults: and maxima: commented out. An uncommented value is a real ceiling, and Kraft must not silently cap someone who never chose a number. A safety field listed only in maxima, such as allowed_tools, starts at its maximum and can only be narrowed. That is why the reference example sets no maxima.allowed_tools: a list there binds every agent task on every harness. gemini cannot enforce a tool list and refuses to launch under one. codex and cursor enforce it through Kraft's permission hook, but refuse a list that leaves out the web tools the hook never sees. On claude, the harness every shipped task uses, it removes every built-in tool the list does not name.

Copyright © 2026