Repository navigation
fix(sq-cost): include Drill pipeline cost and provider-qualified models in task reports - #256
Merged
Merged
Conversation
added 9 commits
October 9, 2026 10:51
Contributor
Author
Coding agent usage on this pull request
Token and model breakdown
Source: Pi session JSONL usage records, covering the task lifetime from 2026-10-09T13:29:46.582Z through report generation. Costs are provider-recorded where available, otherwise list-price estimates; subscription usage is not represented as spend. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Fix the per-task cost report published on a task pull request so it covers the full feature cost, including every Drill validation-pipeline invocation. The old report showed a raw floating-point value, mislabeled provider-recorded cost as estimated, hid the model behind the harness label Pi, and omitted pipeline invocations, understating cost and hiding the models used for correction. Extend only bin/sq-cost.sh and bin/sq-cost-lib.sh; do not change Drill, its schema or recording behavior, and do not create a parallel cost subsystem. Read Drill records from ~/.drill/state.sqlite with SQUAD_DRILL_STATE override, joining agent_invocations to runs by run_id and attributing only the task branch sq/; expose per-invocation agent, model, model_provider, step_name, started_at, and token columns. Price invocations using the existing pricing path: prefer provider-recorded amounts when present, otherwise list-price estimates; OpenCode Go flat-rate subscription usage must never be presented as spend. Merge operator Pi session JSONL and pipeline invocations into provider/model totals, including per-model and overall token and cost totals. Render all Markdown and JSON money to cents, label provider-recorded versus estimated amounts, and add a provider-qualified Models column while retaining the per-model breakdown. Keep the existing publish path and client-visible-project guard exactly unchanged. Preserve no-Drill-run behavior and real nonzero Pi reporting. Add focused regression tests for SQLite attribution, non-operator pipeline models, combined totals, cents formatting, and flat-rate not-spend behavior; update the command documentation and affected script headers. Acceptance: report includes provider-qualified operator and pipeline models and totals; no raw floats appear in report JSON or Markdown; subscription usage is not spend; existing no-run and Pi behavior remain covered; the regression fails on the previous omission. Out of scope: no Drill/schema/recording changes, no new cost subsystem, no unrelated refactors, and no change to publication or the client-visible guard. Commander decisions: provider attribution uses the message-level provider when present, falling back to the preceding model_change; total label says provider-recorded + estimate only when both cost bases exist; provider opencode-go with unprefixed model mimo-v2.5 is flat-rate not-spend while its tokens count. Verification before this run: bash tests/sq-cost.test.sh, bash bin/sq-lint.sh, and bash bin/sq-doc-audience-check.sh all pass. Include this real report output in the PR body as verification evidence: bin/sq-cost.sh report squad-plan-execute-dispatch-contract --json returned: {"found":true,"task":"squad-plan-execute-dispatch-contract","agent":"drill","sessions":0,"started":"1791501217","models":[{"model":"unknown","provider":"unknown","sessions":0,"invocations":7,"input":399939,"output":139517,"cache_read":9916160,"cache_write":0,"estimate_input":399939,"estimate_output":139517,"estimate_cache_read":9916160,"estimate_cache_write":0,"total":10455616,"reported_cost":null,"cost":"$6.27","cost_basis":"estimate","estimate_cost":"$6.27"}],"reason":"task metadata is unavailable","total_cost":{"combined":"$6.27","provider_recorded":null,"estimate":"$6.27","flat_rate_subscription":null}}
What Changed
agent_invocationstorunsbyrun_idin~/.drill/state.sqlite(overridable withSQUAD_DRILL_STATE) and attribute records only to the task's exactsq/<task-id>branch, exposing per-invocation agent, provider/model, step, timestamp, and token columns.cost_basisofprovider-recorded,estimate, orprovider-recorded + estimate), and keep OpenCode Go flat-rate subscription usage out of spend.Verification evidence:
bin/sq-cost.sh report squad-plan-execute-dispatch-contract --jsonreturned:{"found":true,"task":"squad-plan-execute-dispatch-contract","agent":"drill","sessions":0,"started":"1791501217","models":[{"model":"unknown","provider":"unknown","sessions":0,"invocations":7,"input":399939,"output":139517,"cache_read":9916160,"cache_write":0,"estimate_input":399939,"estimate_output":139517,"estimate_cache_read":9916160,"estimate_cache_write":0,"total":10455616,"reported_cost":null,"cost":"$6.27","cost_basis":"estimate","estimate_cost":"$6.27"}],"reason":"task metadata is unavailable","total_cost":{"combined":"$6.27","provider_recorded":null,"estimate":"$6.27","flat_rate_subscription":null}}Risk Assessment
✅ Low: The change is well-bounded, the prior trap regression and test non-hermeticity are correctly fixed and covered by executable regression assertions, and the SQLite query matches the real Drill schema with no reachable correctness gaps found.
Testing
I exercised the real
bin/sq-cost.sh reportCLI against an actual Drill state database and against a controlled operator+Drill fixture, comparing base and target revisions, and ran the focusedtests/sq-cost.test.shsuite. The real report for squad-plan-execute-dispatch-contract matches the intent-provided JSON exactly; the target fixture report shows provider-qualified operator and pipeline models (openai-codex/gpt-6-luna, opencode-go/mimo-v2.5), combined totals in cents with separate provider-recorded/estimate labels, all pipeline invocations included, subscription usage rendered as 'not spend' while its tokens still count, and no raw floats in either JSON or Markdown. The same fixture on base commit 080f6cb drops the pipeline invocations and prints raw$0.1257902, and the new test suite fails on base, confirming the regression catches the previous omission. No actionable failures were found.Evidence: Real report JSON for squad-plan-execute-dispatch-contract (matches intent-provided expected output exactly)
Evidence: Target JSON report: provider-qualified operator+pipeline models, combined cents totals, flat-rate not spend
{"total_cost":{"combined":"$0.80","provider_recorded":"$0.13","estimate":"$0.68","flat_rate_subscription":"not spend"}}Evidence: Target Markdown report with qualified Models column and labeled cents costs
Evidence: Base JSON report showing the previous omission: only operator row, no pipeline invocations
Evidence: Base Markdown report showing raw float $0.1257902 and omitted pipeline cost
Evidence: New regression suite run against base commit 080f6cb (fails on previous omission, exit 1)
Evidence: Base-vs-target CLI harness output
Pipeline
Updates from git push drill
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed ✅
bin/sq-cost.sh:149- The RETURN traptrap 'rm -f "$enriched"' RETURNleaks out ofcmd_reportinto its caller.cmd_reportis invoked directly at top level forreport(line 391) and via a subshell forpublish(line 262), so the trap fires once whileenrichedis still in scope there. Butcmd_taskcallscmd_report "$task_id" --jsonin the same shell (line 288), so the trap fires a second time oncmd_task's return, after thecmd_reportlocal is out of scope; underset -uthis aborts withbin/sq-cost.sh: line 288: enriched: unbound variableand exit 1, even though the JSON was already written. At base,--jsonreturned before the trap was set, so this is a regression. Reproduction (any found=true task, Pi-only or Drill):bash bin/sq-cost.sh task <id> --jsonprints the report then exits 1;bash bin/sq-cost.sh report <id> --jsonexits 0. Impact:task <id> --jsonis the documented equivalent ofreport --json(script header) and is the exact callbin/sq-trajectory.shmakes ("$COST" task "$id" --json ... || cost_json='{}'), so callers that check status silently lose the report. No existing test executestask <id> --jsonwith found=true, so the suite does not catch it. Fix: make the cleanup self-clearing and unset-safe, e.g.trap 'rm -f "${enriched:-}"; trap - RETURN' RETURN, and add an executable regression assertingtask <id> --jsonexits 0.bin/sq-cost.sh:111-cmd_reportnow always reads$HOME/.drill/state.sqlitewhen the Drill file exists, soreport/publishand anything that calls them transitively depend on the developer's real Drill state unlessSQUAD_DRILL_STATEis set.tests/sq-cost.test.shhandles this with a global export, buttests/sq-trajectory.test.shinvokessq-cost.sh task <id> --jsonwithout overriding it; if a real run exists on branchsq/alpha(orsq/beta/sq/taskNN), the report becomes found=true and the test's.tokens.input.value==nullcontract breaks. This is currently masked by the trap bug above (that non-zero exit forcescost_json='{}'), so fixing the trap will expose the non-hermetic read. Recommend sandboxingSQUAD_DRILL_STATEin that suite (or in the test runner) alongside the fix.🔧 Fix: fix(sq-cost): self-clear report trap and isolate test Drill state
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/sq-cost.test.sh(all focused behavior tests pass, including Drill attribution, non-operator pipeline models, combined totals, cents formatting, flat-rate not-spend, message-level provider attribution, and Drill read-failure visibility)bin/sq-cost.sh report squad-plan-execute-dispatch-contract --jsonagainst the real~/.drill/state.sqliteand compared withjq -Sto the intent-provided expected output (exact semantic match)jq '[paths(type=="number") as $p | select($p[-1]|test("cost"))]'over target and real report JSON to confirm no numeric money fields remainStandalone CLI harness building a synthetic operator Pi session plus Drill SQLite fixture, then runningsq-cost.sh report pipeline-task(JSON and Markdown) from both base commit080f6cband targetde9badfbash tests/sq-cost.test.shagainst agit archiveextraction of base commit080f6cbto confirm the new regression fails on the previous omission (exit 1 on the cents assertion)git status --porcelainto confirm no residual worktree changes after testing✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.