Result file
Every agent and subprocess task gets KRAFT_RESULT_PATH in its environment:
the path of a JSON file it writes before it exits. Kraft reads the task's
status, findings, verdict and usage back from it. The file is
run/results/<session id>.json under $KRAFT_HOME, and it stays after the
session ends.
Who must write it
- An agent task must. Kraft's instructions to every agent say what to
write. An agent that exits cleanly without a result file ends
failed, and so does one whose turn ends with a background job still running. - A subprocess task may. Without one, its exit code decides. See Subprocess tasks.
Example
{
"status": "done",
"session_summary_ref": ".engineering/sessions/2c1f.md",
"findings": [
{
"severity": "important",
"message": "The retry loop never backs off, so a down server gets hammered.",
"file": "src/client.py",
"line": 42,
"source_plugin": "code-review"
}
]
}
Fields
The file is one JSON object. Kraft ignores keys it does not know.
| Field | Type | Read from | Meaning |
|---|---|---|---|
status | string | every task | How the task went. See status. |
concerns | string | every task | With done_with_concerns, why. From a judge or a gate review, the reasoning. |
question | string | every task | With needs_context, what the task needs to know. The item stops with it. |
session_summary_ref | string | every task | The session summary's path, relative to the repo root. |
findings | array | measuring tasks | What a review or check found. See findings. |
verdict | string | a fix-loop judge, a gate's auto_review | See verdict. |
suggested_action | object | the last task a node ran before it stopped | {"action": "skip" | "retry" | "abandon", "reason": "..."}. kraft view show prints the one command that takes it. Any other action is ignored. |
usage, cost_usd, model | object, number, string | every task | See usage. |
status
| Value | Meaning |
|---|---|
done | The task finished. |
done_with_concerns | The task finished but doubts the result; concerns says why. The node goes on. A repair in an on_failure recovery that ends this way stops the item for you. |
failed | The task could not finish. |
needs_context | The task is missing information; question says what. The item stops for you with the question. Ask for everything at once: each question costs a new run. |
A file that is not valid JSON, is not UTF-8, has no status or has any
other value ends the task failed. An empty file counts as no file.
findings
Each entry is an object:
| Key | Required | Meaning |
|---|---|---|
severity | yes | critical, important or minor. Which ones open a fix loop is findings.loop_severities in policy.yaml. |
message | yes | What is wrong and why. |
source_plugin | yes | Who found it, such as code-review. |
file | no | The path, relative to the repo root. |
line | no | An integer. Not part of the finding's identity, so an approximate line is fine. |
same_as | no | The 16-character tag of a finding from an earlier round that this one restates. Only a tag the task was shown counts; any other is ignored. |
An entry missing a required key, or with an unknown severity, is dropped. The
findings_measured event counts dropped entries under dropped, per task.
A finding's identity across rounds is its same_as, or else a hash of
source_plugin, file and message. The fix loop's stuck check and the
judge's history read it. See Fix loop and judge.
verdict
A verdict from a task that ended failed or needs_context does not count,
whatever it says: Kraft reads it as the last column below.
| Written by | Values | Anything else |
|---|---|---|
A fix loop's judge | continue, stop_needs_human, stop_downgrade. See The judge. | continue |
A gate's auto_review | approve, reject, fixed (it repaired the worktree and the node runs again), undecided. reject and fixed need concerns, which becomes the steer for the re-run. | undecided: the gate waits for you |
usage
Kraft reads a session's tokens and cost from the result file first, and from the agent's log after that when the harness declares a log format. A harness that reports usage only through the result file is asked to write it there.
| Key | Meaning |
|---|---|
usage.input_tokens, usage.output_tokens | Tokens used. |
usage.cache_creation_input_tokens, usage.cache_read_input_tokens | Cache tokens, counted as billed input. |
cost_usd (or total_cost_usd) | What the session cost, at the top level of the file. Leave it out rather than guess. |
model | The model that did the work. |