Fallback and escalation
A task can name where its launch goes next, and an escalation turn runs on a harness you choose.
Fallback
A task that must not wait out a rate limit can name where its launch goes
next. A fallback: list is ordered. Each entry may name a harness: (a
harness profile id from harnesses.yaml) plus one route, the same either-or a
task obeys: an agent profile:, or a model: and/or effort:. It keeps
whatever it omits from the task's own launch.
The list lives on an agent profile, as the tier's default, and a task may set
its own, which replaces the profile's whole (fallback: [] means none, even
under a profile that has one). For example:
# harnesses.yaml
profiles:
deep:
effort: high
model: { claude: opus, codex: gpt-5.6-sol }
fallback:
- { harness: codex } # same tier, other harness
- { profile: strong } # other tier, same harness
strong:
effort: high
model: { claude: sonnet, codex: gpt-5.6-terra }
# library.yaml
tasks:
legacy_task:
kind: agent
harness: claude
model: opus
effort: high
fallback:
- { model: sonnet } # claude, sonnet, high
- { harness: codex, model: gpt-5.6-terra } # codex, gpt-5.6-terra, high
| Entry | on claude + profile: deep | on claude + model: opus, effort: high |
|---|---|---|
{ harness: codex } | codex + deep: gpt-5.6-sol, high | codex + opus: refused, codex takes no opus |
{ profile: strong } | claude + strong: sonnet, high | claude + strong |
{ model: sonnet } | claude + sonnet, no effort | claude + sonnet, high |
An entry's own profile's list is not followed: lists do not chain. Like the
profile body, a profile's list is read live at each launch. Through extends,
a child task's fallback: replaces its parent's.
Every entry of every list a chain's task uses must pair with its harness, the
same check a task's own route gets: in Settings → Harnesses (which shows each
profile's list with its problems per entry), in the guard on a harness save,
and in kraft admin doctor. A problem names the task, the list's source and
the entry's index: chain 'default' task 'implementation.main.implementer': fallback entry 0 (profile 'deep''s list): .... A disabled harness is not a
pairing problem, since the launch skips it.
The candidates are the task's own launch, then each entry in order. Kraft moves to the next candidate, in the same dispatch, when:
- a launch ends rate-limited (a harness that declares
rate_limit_signalreported a rejected request). The next launch is a fresh session in the same worktree, with the task's instruction and a note that the earlier attempt may have left partial work (git status,git diff); - a candidate is unavailable before anything launches: its harness
profile is missing, disabled or sets a default Kraft cannot apply, its agent
profile is missing or has no model for its provider, or its executable is
not on
PATH(not checked under a sandbox, where the executable lives in the container). This is checked again at every launch, so fixing it takes effect at once; - a candidate is known to be limited: Kraft remembers, from the event log,
which harness and model a
rate_limit_hitlimited, on any work item, and skips that pair until its reset. One account-wide limit therefore costs one quickly-refused launch per model before each is remembered.
When no candidate is left and one was limited, the item parks as
rate_limited until the earliest reset among them, and its relaunch starts
from the top of the list, so the first choice is used again as soon as it is
back. When every candidate is unavailable, the task stops for a human with a
reason naming each one and why.
What carries over and what does not:
- The work item's
set-overridesvalues and a node'smodel/effortoverride pick the task's own launch only, and so does a fix loop's escalation model. A fallback runs exactly as its entry says. A node'sextra_promptis part of the instruction and does reach it. - Switching spends none of
rate_limit_retries, which still counts parks. The spend caps inbudget:are checked before every launch, fallbacks included, and a task's time cap covers all of its attempts together. - Every entry's harness must be in the task's
allowed_harnesses. A task's own list is checked when the chain materializes; a profile's list, which is live, is checked at the launch, which skips a refused entry as unavailable.deny_tools,allowed_toolsand the item's sandbox apply to a fallback as to the task's own launch, so a sandboxed fallback onto Cursor or Codex with a tool policy stops as a configuration error rather than skipping to the next entry.
It is opt-in. A task with no list (none of its own, and none on its profile,
or fallback: []) launches, parks and stops exactly as it would without it,
and never consults the memory. No shipped task or profile declares one. A gate's auto_review task cannot declare a fallback: list of its own, and is
refused the same way — at its own launch, naming the gate and the profile —
when its profile: carries one: a gate review launches once and never walks
either list.
Only a harness that declares rate_limit_signal can trigger a switch on a
rate limit, which is claude, codex, opencode and antigravity. gemini, amp
and cursor can be fallback targets, and are skipped when unavailable, but a
rate limit on them fails the launch as it does without a list.
Every skip or switch is logged:
- one
launch_fallbackevent (kraft view events --type launch_fallback), withfrom,to(nullwhen nothing was left),reason(rate_limit_hit,known_limitedorunavailable, with adetail),resets_at_isoandoverride_not_carried; - one sentence on the item's timeline, such as "Ran on claude / sonnet instead of claude / opus: claude / opus is rate-limited until 15:40.";
- one INFO line in the server log;
- a "fallback" marker on the board card while the item's latest launch is a fallback's, with the same sentence as its tooltip.
Session rows record the model that actually ran, so Analytics attributes each attempt's cost to it.
Escalation turns
An escalation turn, the conversation Kraft opens when a person escalates a
stopped item or auto_escalate_stuck does, runs on the harnesses.yaml
profile policy's escalation_harness names: claude unless you set it. Set
it in policy.yaml's defaults:, on a repository, chain or node policy:,
or for one item with kraft item set-policy --policy escalation_harness=codex.
item follows the harness the item's own latest agent task ran on. The
thread id the next turn resumes is read with that harness's own log reader,
and a resumed turn records only what it added to the thread. See
escalation_harness. A gate's automated review is not
this: its auto_review task names its own harness: in the chain.