nextv1.4.0
Agent harnesses

Fallback and escalation

Where a launch goes next when a harness is limited or unavailable, and which harness runs an escalation turn.

A task can name where its launch goes next, and an escalation turn runs on a harness you choose.

Fallback

A task that must not wait out a rate limit can name where its launch goes next. A fallback: list is ordered. Each entry may name a harness: (a harness profile id from harnesses.yaml) plus one route, the same either-or a task obeys: an agent profile:, or a model: and/or effort:. It keeps whatever it omits from the task's own launch.

The list lives on an agent profile, as the tier's default, and a task may set its own, which replaces the profile's whole (fallback: [] means none, even under a profile that has one). For example:

# harnesses.yaml
profiles:
  deep:
    effort: high
    model: { claude: opus, codex: gpt-5.6-sol }
    fallback:
      - { harness: codex }            # same tier, other harness
      - { profile: strong }           # other tier, same harness
  strong:
    effort: high
    model: { claude: sonnet, codex: gpt-5.6-terra }
# library.yaml
tasks:
  legacy_task:
    kind: agent
    harness: claude
    model: opus
    effort: high
    fallback:
      - { model: sonnet }                          # claude, sonnet, high
      - { harness: codex, model: gpt-5.6-terra }   # codex, gpt-5.6-terra, high
Entryon claude + profile: deepon claude + model: opus, effort: high
{ harness: codex }codex + deep: gpt-5.6-sol, highcodex + opus: refused, codex takes no opus
{ profile: strong }claude + strong: sonnet, highclaude + strong
{ model: sonnet }claude + sonnet, no effortclaude + sonnet, high

An entry's own profile's list is not followed: lists do not chain. Like the profile body, a profile's list is read live at each launch. Through extends, a child task's fallback: replaces its parent's.

Every entry of every list a chain's task uses must pair with its harness, the same check a task's own route gets: in Settings → Harnesses (which shows each profile's list with its problems per entry), in the guard on a harness save, and in kraft admin doctor. A problem names the task, the list's source and the entry's index: chain 'default' task 'implementation.main.implementer': fallback entry 0 (profile 'deep''s list): .... A disabled harness is not a pairing problem, since the launch skips it.

The candidates are the task's own launch, then each entry in order. Kraft moves to the next candidate, in the same dispatch, when:

  • a launch ends rate-limited (a harness that declares rate_limit_signal reported a rejected request). The next launch is a fresh session in the same worktree, with the task's instruction and a note that the earlier attempt may have left partial work (git status, git diff);
  • a candidate is unavailable before anything launches: its harness profile is missing, disabled or sets a default Kraft cannot apply, its agent profile is missing or has no model for its provider, or its executable is not on PATH (not checked under a sandbox, where the executable lives in the container). This is checked again at every launch, so fixing it takes effect at once;
  • a candidate is known to be limited: Kraft remembers, from the event log, which harness and model a rate_limit_hit limited, on any work item, and skips that pair until its reset. One account-wide limit therefore costs one quickly-refused launch per model before each is remembered.

When no candidate is left and one was limited, the item parks as rate_limited until the earliest reset among them, and its relaunch starts from the top of the list, so the first choice is used again as soon as it is back. When every candidate is unavailable, the task stops for a human with a reason naming each one and why.

What carries over and what does not:

  • The work item's set-overrides values and a node's model/effort override pick the task's own launch only, and so does a fix loop's escalation model. A fallback runs exactly as its entry says. A node's extra_prompt is part of the instruction and does reach it.
  • Switching spends none of rate_limit_retries, which still counts parks. The spend caps in budget: are checked before every launch, fallbacks included, and a task's time cap covers all of its attempts together.
  • Every entry's harness must be in the task's allowed_harnesses. A task's own list is checked when the chain materializes; a profile's list, which is live, is checked at the launch, which skips a refused entry as unavailable. deny_tools, allowed_tools and the item's sandbox apply to a fallback as to the task's own launch, so a sandboxed fallback onto Cursor or Codex with a tool policy stops as a configuration error rather than skipping to the next entry.

It is opt-in. A task with no list (none of its own, and none on its profile, or fallback: []) launches, parks and stops exactly as it would without it, and never consults the memory. No shipped task or profile declares one. A gate's auto_review task cannot declare a fallback: list of its own, and is refused the same way — at its own launch, naming the gate and the profile — when its profile: carries one: a gate review launches once and never walks either list.

Only a harness that declares rate_limit_signal can trigger a switch on a rate limit, which is claude, codex, opencode and antigravity. gemini, amp and cursor can be fallback targets, and are skipped when unavailable, but a rate limit on them fails the launch as it does without a list.

Every skip or switch is logged:

  • one launch_fallback event (kraft view events --type launch_fallback), with from, to (null when nothing was left), reason (rate_limit_hit, known_limited or unavailable, with a detail), resets_at_iso and override_not_carried;
  • one sentence on the item's timeline, such as "Ran on claude / sonnet instead of claude / opus: claude / opus is rate-limited until 15:40.";
  • one INFO line in the server log;
  • a "fallback" marker on the board card while the item's latest launch is a fallback's, with the same sentence as its tooltip.

Session rows record the model that actually ran, so Analytics attributes each attempt's cost to it.

Escalation turns

An escalation turn, the conversation Kraft opens when a person escalates a stopped item or auto_escalate_stuck does, runs on the harnesses.yaml profile policy's escalation_harness names: claude unless you set it. Set it in policy.yaml's defaults:, on a repository, chain or node policy:, or for one item with kraft item set-policy --policy escalation_harness=codex. item follows the harness the item's own latest agent task ran on. The thread id the next turn resumes is read with that harness's own log reader, and a resumed turn records only what it added to the thread. See escalation_harness. A gate's automated review is not this: its auto_review task names its own harness: in the chain.

Copyright © 2026