fix(gatorwalk-factory): record_usage takes the token total a harness reports (swamp-club #2778) #393

Merged
seth merged 1 commit from cue/2778-gatorwalk-factory-record into main 2026-09-30 21:43:11 +00:00
Owner

Fixes swamp-club #2778.

Claude Code reports one subagent_tokens total per subagent, with no input/output split, so record_usage (which required both halves) could not take it: the cue trial ended with every dispatch "without usage".

Engine

  • A usage has totalTokens, or inputTokens and outputTokens together, or all three; toolUses and durationMs are optional. The total is not checked against the split (a harness total can count cache tokens). Half a split, or no tokens, is refused with a reason. All run-record changes are optional fields, so existing records stay valid.
  • Each dispatch records its stage's work mode. Metrics sum the total (a dispatch's totalTokens, else input + output), keep the split over the dispatches that gave one, sum tool uses and duration, and count dispatches without usage by mode (older dispatches fall back to the pinned definition's stage mode).
  • The summary's token line leads with the total, then the split, tool uses and time, then "N without usage: …" by mode, then per-model totals.

Skill (usage sections only)

  • The count comes from the harness's report of the subagent (the task notification's <usage> block, or the Agent tool result when synchronous), never from the subagent; don't ask reviewers for tokens; interactive dispatches normally have none.
  • The worked example records totalTokens, and skill_test.ts checks that the attested total equals the sum of the example's reported counts.

Verification: verify-build e0a185d6 and verify-reviews 5ac1d7a3 on 91e7b3c5, all steps passed (2 guarded skips); both reviews pass with low findings only (per-model split has no per-model split count; new optional fields are not readable by an older strict engine; an empty coerced token input becomes 0). Attestation d1f623ce.

🤖 Generated with Claude Code

Fixes swamp-club #2778. Claude Code reports one `subagent_tokens` total per subagent, with no input/output split, so `record_usage` (which required both halves) could not take it: the cue trial ended with every dispatch "without usage". **Engine** - A usage has `totalTokens`, or `inputTokens` and `outputTokens` together, or all three; `toolUses` and `durationMs` are optional. The total is not checked against the split (a harness total can count cache tokens). Half a split, or no tokens, is refused with a reason. All run-record changes are optional fields, so existing records stay valid. - Each dispatch records its stage's work mode. Metrics sum the total (a dispatch's `totalTokens`, else input + output), keep the split over the dispatches that gave one, sum tool uses and duration, and count dispatches without usage by mode (older dispatches fall back to the pinned definition's stage mode). - The summary's token line leads with the total, then the split, tool uses and time, then "N without usage: …" by mode, then per-model totals. **Skill** (usage sections only) - The count comes from the harness's report of the subagent (the task notification's `<usage>` block, or the Agent tool result when synchronous), never from the subagent; don't ask reviewers for tokens; interactive dispatches normally have none. - The worked example records `totalTokens`, and `skill_test.ts` checks that the attested total equals the sum of the example's reported counts. **Verification**: verify-build `e0a185d6` and verify-reviews `5ac1d7a3` on 91e7b3c5, all steps passed (2 guarded skips); both reviews pass with low findings only (per-model split has no per-model split count; new optional fields are not readable by an older strict engine; an empty coerced token input becomes 0). Attestation `d1f623ce`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(gatorwalk-factory): record_usage takes the token total a harness reports (swamp-club #2778)
All checks were successful
CI / Review Integrity (pull_request) Successful in 46s
CI / Validate Attestation (pull_request) Successful in 48s
91e7b3c584
Claude Code reports one subagent_tokens total per subagent, with no
input/output split, so record_usage (which required both halves) could
not take it and every dispatch ended without usage.

- A usage has totalTokens, or inputTokens and outputTokens together, or
  all three; toolUses and durationMs are optional. The total is not
  checked against the split.
- Each dispatch records its stage's work mode. Metrics sum the total,
  keep the split over the dispatches that gave one, and count dispatches
  without usage by mode; the summary shows all of it.
- The skill says the count comes from the harness's report of the
  subagent (the task notification's usage block), never from the
  subagent, and that interactive dispatches normally have none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seth merged commit 48bb715fda into main 2026-09-30 21:43:11 +00:00
seth deleted branch cue/2778-gatorwalk-factory-record 2026-09-30 21:43:12 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
swamp-club/swamp-extensions!393
No description provided.