Skip to content

fix(sq-cost): include Drill pipeline cost and provider-qualified models in task reports - #256

Merged
90sRehem merged 9 commits into
mainfrom
sq/squad-cost-report-models-and-pipeline
Oct 9, 2026
Merged

90sRehem merged 9 commits into
mainfrom
sq/squad-cost-report-models-and-pipeline

Conversation

@90sRehem

@90sRehem 90sRehem commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Intent

Fix the per-task cost report published on a task pull request so it covers the full feature cost, including every Drill validation-pipeline invocation. The old report showed a raw floating-point value, mislabeled provider-recorded cost as estimated, hid the model behind the harness label Pi, and omitted pipeline invocations, understating cost and hiding the models used for correction. Extend only bin/sq-cost.sh and bin/sq-cost-lib.sh; do not change Drill, its schema or recording behavior, and do not create a parallel cost subsystem. Read Drill records from ~/.drill/state.sqlite with SQUAD_DRILL_STATE override, joining agent_invocations to runs by run_id and attributing only the task branch sq/; expose per-invocation agent, model, model_provider, step_name, started_at, and token columns. Price invocations using the existing pricing path: prefer provider-recorded amounts when present, otherwise list-price estimates; OpenCode Go flat-rate subscription usage must never be presented as spend. Merge operator Pi session JSONL and pipeline invocations into provider/model totals, including per-model and overall token and cost totals. Render all Markdown and JSON money to cents, label provider-recorded versus estimated amounts, and add a provider-qualified Models column while retaining the per-model breakdown. Keep the existing publish path and client-visible-project guard exactly unchanged. Preserve no-Drill-run behavior and real nonzero Pi reporting. Add focused regression tests for SQLite attribution, non-operator pipeline models, combined totals, cents formatting, and flat-rate not-spend behavior; update the command documentation and affected script headers. Acceptance: report includes provider-qualified operator and pipeline models and totals; no raw floats appear in report JSON or Markdown; subscription usage is not spend; existing no-run and Pi behavior remain covered; the regression fails on the previous omission. Out of scope: no Drill/schema/recording changes, no new cost subsystem, no unrelated refactors, and no change to publication or the client-visible guard. Commander decisions: provider attribution uses the message-level provider when present, falling back to the preceding model_change; total label says provider-recorded + estimate only when both cost bases exist; provider opencode-go with unprefixed model mimo-v2.5 is flat-rate not-spend while its tokens count. Verification before this run: bash tests/sq-cost.test.sh, bash bin/sq-lint.sh, and bash bin/sq-doc-audience-check.sh all pass. Include this real report output in the PR body as verification evidence: bin/sq-cost.sh report squad-plan-execute-dispatch-contract --json returned: {"found":true,"task":"squad-plan-execute-dispatch-contract","agent":"drill","sessions":0,"started":"1791501217","models":[{"model":"unknown","provider":"unknown","sessions":0,"invocations":7,"input":399939,"output":139517,"cache_read":9916160,"cache_write":0,"estimate_input":399939,"estimate_output":139517,"estimate_cache_read":9916160,"estimate_cache_write":0,"total":10455616,"reported_cost":null,"cost":"$6.27","cost_basis":"estimate","estimate_cost":"$6.27"}],"reason":"task metadata is unavailable","total_cost":{"combined":"$6.27","provider_recorded":null,"estimate":"$6.27","flat_rate_subscription":null}}

What Changed

  • Join agent_invocations to runs by run_id in ~/.drill/state.sqlite (overridable with SQUAD_DRILL_STATE) and attribute records only to the task's exact sq/<task-id> branch, exposing per-invocation agent, provider/model, step, timestamp, and token columns.
  • Merge Drill pipeline invocations with attributable operator Pi session rows into provider/model totals, and render provider-qualified model labels alongside the per-model token and cost breakdown in Markdown and JSON.
  • Format all JSON and Markdown money to cents, label provider-recorded versus estimate amounts (cost_basis of provider-recorded, estimate, or provider-recorded + estimate), and keep OpenCode Go flat-rate subscription usage out of spend.

Verification evidence: bin/sq-cost.sh report squad-plan-execute-dispatch-contract --json returned:

{"found":true,"task":"squad-plan-execute-dispatch-contract","agent":"drill","sessions":0,"started":"1791501217","models":[{"model":"unknown","provider":"unknown","sessions":0,"invocations":7,"input":399939,"output":139517,"cache_read":9916160,"cache_write":0,"estimate_input":399939,"estimate_output":139517,"estimate_cache_read":9916160,"estimate_cache_write":0,"total":10455616,"reported_cost":null,"cost":"$6.27","cost_basis":"estimate","estimate_cost":"$6.27"}],"reason":"task metadata is unavailable","total_cost":{"combined":"$6.27","provider_recorded":null,"estimate":"$6.27","flat_rate_subscription":null}}

Risk Assessment

✅ Low: The change is well-bounded, the prior trap regression and test non-hermeticity are correctly fixed and covered by executable regression assertions, and the SQLite query matches the real Drill schema with no reachable correctness gaps found.

Testing

I exercised the real bin/sq-cost.sh report CLI against an actual Drill state database and against a controlled operator+Drill fixture, comparing base and target revisions, and ran the focused tests/sq-cost.test.sh suite. The real report for squad-plan-execute-dispatch-contract matches the intent-provided JSON exactly; the target fixture report shows provider-qualified operator and pipeline models (openai-codex/gpt-6-luna, opencode-go/mimo-v2.5), combined totals in cents with separate provider-recorded/estimate labels, all pipeline invocations included, subscription usage rendered as 'not spend' while its tokens still count, and no raw floats in either JSON or Markdown. The same fixture on base commit 080f6cb drops the pipeline invocations and prints raw $0.1257902, and the new test suite fails on base, confirming the regression catches the previous omission. No actionable failures were found.

Evidence: Real report JSON for squad-plan-execute-dispatch-contract (matches intent-provided expected output exactly)
{
  "found": true,
  "task": "squad-plan-execute-dispatch-contract",
  "agent": "drill",
  "sessions": 0,
  "started": "1791501217",
  "models": [
    {
      "model": "unknown",
      "provider": "unknown",
      "sessions": 0,
      "invocations": 7,
      "input": 399939,
      "output": 139517,
      "cache_read": 9916160,
      "cache_write": 0,
      "estimate_input": 399939,
      "estimate_output": 139517,
      "estimate_cache_read": 9916160,
      "estimate_cache_write": 0,
      "total": 10455616,
      "reported_cost": null,
      "cost": "$6.27",
      "cost_basis": "estimate",
      "estimate_cost": "$6.27"
    }
  ],
  "reason": "task metadata is unavailable",
  "total_cost": {
    "combined": "$6.27",
    "provider_recorded": null,
    "estimate": "$6.27",
    "flat_rate_subscription": null
  }
}
Evidence: Target JSON report: provider-qualified operator+pipeline models, combined cents totals, flat-rate not spend

{"total_cost":{"combined":"$0.80","provider_recorded":"$0.13","estimate":"$0.68","flat_rate_subscription":"not spend"}}

{
  "found": true,
  "task": "pipeline-task",
  "agent": "pi",
  "worktree": "/tmp/sq-cost-evidence.0zoWaP/worktree",
  "sessions": 1,
  "started": "2026-01-01T00:00:00Z",
  "models": [
    {
      "model": "gpt-6-luna",
      "provider": "openai-codex",
      "sessions": 1,
      "invocations": 2,
      "input": 150100,
      "output": 15050,
      "cache_read": 0,
      "cache_write": 0,
      "estimate_input": 150000,
      "estimate_output": 15000,
      "estimate_cache_read": 0,
      "estimate_cache_write": 0,
      "total": 165150,
      "reported_cost": "$0.13",
      "cost": "$0.80",
      "cost_basis": "provider-recorded + estimate",
      "estimate_cost": "$0.68"
    },
    {
      "model": "mimo-v2.5",
      "provider": "opencode-go",
      "sessions": 0,
      "invocations": 1,
      "input": 1000,
      "output": 500,
      "cache_read": 0,
      "cache_write": 0,
      "estimate_input": 1000,
      "estimate_output": 500,
      "estimate_cache_read": 0,
      "estimate_cache_write": 0,
      "total": 1500,
      "reported_cost": null,
      "cost": null,
      "cost_basis": "flat-rate subscription",
      "estimate_cost": null
    }
  ],
  "total_cost": {
    "combined": "$0.80",
    "provider_recorded": "$0.13",
    "estimate": "$0.68",
    "flat_rate_subscription": "not spend"
  }
}
Evidence: Target Markdown report with qualified Models column and labeled cents costs
## Coding agent usage on this pull request

| Contributor | Agent | Models | Sessions | Total tokens | Cost |
|---|---|---|---:|---:|---|
| Squad task pipeline-task | Pi | openai-codex/gpt-6-luna, opencode-go/mimo-v2.5 | 1 | 166.7 thousand | provider-recorded + estimate total: $0.80; provider-recorded: $0.13; estimate: $0.68; flat-rate subscription: not spend |

### Token and model breakdown

| Model | Input | Output | Cache read | Cache write | Total tokens | Cost |
|---|---:|---:|---:|---:|---:|---|
| openai-codex/gpt-6-luna | 150.1 thousand | 15.1 thousand | 0 | 0 | 165.2 thousand | provider-recorded: $0.13; estimate: $0.68 |
| opencode-go/mimo-v2.5 | 1 thousand | 500 | 0 | 0 | 1.5 thousand | flat-rate subscription: not spend |

_Source: Pi session JSONL usage records and Drill agent-invocation records, covering the task lifetime through report generation. Provider-recorded costs and list-price estimates are labeled separately; subscription usage is not represented as spend._
Evidence: Base JSON report showing the previous omission: only operator row, no pipeline invocations
{
  "found": true,
  "task": "pipeline-task",
  "agent": "pi",
  "worktree": "/tmp/sq-cost-evidence.0zoWaP/worktree",
  "sessions": 1,
  "started": "2026-01-01T00:00:00Z",
  "models": [
    {
      "model": "gpt-6-luna",
      "sessions": 1,
      "input": 100,
      "output": 50,
      "cache_read": 0,
      "cache_write": 0,
      "total": 150,
      "reported_cost": 0.1257902,
      "provider": "openai-codex"
    }
  ]
}
Evidence: Base Markdown report showing raw float $0.1257902 and omitted pipeline cost
## Coding agent usage on this pull request

| Contributor | Agent | Sessions | Total tokens | Estimated cost |
|---|---|---:|---:|---:|
| Squad task pipeline-task | Pi | 1 | 150 | $0.1257902 |

### Token and model breakdown

| Model | Input | Output | Cache read | Cache write | Total tokens | Estimated cost |
|---|---:|---:|---:|---:|---:|---:|
| gpt-6-luna | 100 | 50 | 0 | 0 | 150 | provider-recorded: 0.1257902 |

_Source: Pi session JSONL usage records, covering the task lifetime from 2026-01-01T00:00:00Z through report generation. Costs are provider-recorded where available, otherwise list-price estimates; subscription usage is not represented as spend._
Evidence: New regression suite run against base commit 080f6cb (fails on previous omission, exit 1)
ok - opus pricing resolves correctly
ok - sonnet pricing resolves correctly
ok - haiku pricing resolves correctly
ok - gpt-4o pricing resolves correctly
ok - unknown model defaults to sonnet pricing
ok - 1M input + 1M output sonnet = $18
ok - 500K input + 200K output haiku ≈ $1.20
ok - cache tokens included in cost calculation
ok - JSONL transcript parsing extracts correct token counts
ok - end-to-end: synthetic transcript → correct cost
ok - directory scanning sums multiple transcripts
ok - missing transcript returns zero gracefully
ok - empty transcript (no assistant messages) returns zero cost
ok - missing directory returns zero gracefully
ok - no-meta task cost lookup returns estimate line gracefully
ok - transcript on missing file returns gracefully
ok - dir on missing directory returns gracefully
ok - model normalization works correctly
ok - CLI transcript command works
ok - CLI estimate command works
ok - CLI dir command works
ok - CLI price command works
ok - CLI pricing-table command works
not ok - Pi report preserves provider cost across attempts in cents (missing: '"reported_cost": "$0.05"')
--- output ---
{
  "found": true,
  "task": "pi-task",
  "agent": "pi",
  "worktree": "/tmp/sq-cost.VGjRmU/pi-worktree",
  "sessions": 2,
  "started": "2025-12-20T00:00:00Z",
  "models": [
    {
      "model": "claude-sonnet-4",
      "sessions": 2,
      "input": 2600000400,
      "output": 253,
      "cache_read": 7,
      "cache_write": 2,
      "total": 2600000662,
      "reported_cost": 0.05,
      "provider": "anthropic"
    }
  ]
}
Evidence: Base-vs-target CLI harness output
base json exit=0
base md exit=0
target json exit=0
target md exit=0
=== BASE JSON ===
{
  "found": true,
  "task": "pipeline-task",
  "agent": "pi",
  "worktree": "/tmp/sq-cost-evidence.0zoWaP/worktree",
  "sessions": 1,
  "started": "2026-01-01T00:00:00Z",
  "models": [
    {
      "model": "gpt-6-luna",
      "sessions": 1,
      "input": 100,
      "output": 50,
      "cache_read": 0,
      "cache_write": 0,
      "total": 150,
      "reported_cost": 0.1257902,
      "provider": "openai-codex"
    }
  ]
}
=== TARGET JSON ===
{
  "found": true,
  "task": "pipeline-task",
  "agent": "pi",
  "worktree": "/tmp/sq-cost-evidence.0zoWaP/worktree",
  "sessions": 1,
  "started": "2026-01-01T00:00:00Z",
  "models": [
    {
      "model": "gpt-6-luna",
      "provider": "openai-codex",
      "sessions": 1,
      "invocations": 2,
      "input": 150100,
      "output": 15050,
      "cache_read": 0,
      "cache_write": 0,
      "estimate_input": 150000,
      "estimate_output": 15000,
      "estimate_cache_read": 0,
      "estimate_cache_write": 0,
      "total": 165150,
      "reported_cost": "$0.13",
      "cost": "$0.80",
      "cost_basis": "provider-recorded + estimate",
      "estimate_cost": "$0.68"
    },
    {
      "model": "mimo-v2.5",
      "provider": "opencode-go",
      "sessions": 0,
      "invocations": 1,
      "input": 1000,
      "output": 500,
      "cache_read": 0,
      "cache_write": 0,
      "estimate_input": 1000,
      "estimate_output": 500,
      "estimate_cache_read": 0,
      "estimate_cache_write": 0,
      "total": 1500,
      "reported_cost": null,
      "cost": null,
      "cost_basis": "flat-rate subscription",
      "estimate_cost": null
    }
  ],
  "total_cost": {
    "combined": "$0.80",
    "provider_recorded": "$0.13",
    "estimate": "$0.68",
    "flat_rate_subscription": "not spend"
  }
}
=== BASE MARKDOWN ===
## Coding agent usage on this pull request

| Contributor | Agent | Sessions | Total tokens | Estimated cost |
|---|---|---:|---:|---:|
| Squad task pipeline-task | Pi | 1 | 150 | $0.1257902 |

### Token and model breakdown

| Model | Input | Output | Cache read | Cache write | Total tokens | Estimated cost |
|---|---:|---:|---:|---:|---:|---:|
| gpt-6-luna | 100 | 50 | 0 | 0 | 150 | provider-recorded: 0.1257902 |

_Source: Pi session JSONL usage records, covering the task lifetime from 2026-01-01T00:00:00Z through report generation. Costs are provider-recorded where available, otherwise list-price estimates; subscription usage is not represented as spend._
=== TARGET MARKDOWN ===
## Coding agent usage on this pull request

| Contributor | Agent | Models | Sessions | Total tokens | Cost |
|---|---|---|---:|---:|---|
| Squad task pipeline-task | Pi | openai-codex/gpt-6-luna, opencode-go/mimo-v2.5 | 1 | 166.7 thousand | provider-recorded + estimate total: $0.80; provider-recorded: $0.13; estimate: $0.68; flat-rate subscription: not spend |

### Token and model breakdown

| Model | Input | Output | Cache read | Cache write | Total tokens | Cost |
|---|---:|---:|---:|---:|---:|---|
| openai-codex/gpt-6-luna | 150.1 thousand | 15.1 thousand | 0 | 0 | 165.2 thousand | provider-recorded: $0.13; estimate: $0.68 |
| opencode-go/mimo-v2.5 | 1 thousand | 500 | 0 | 0 | 1.5 thousand | flat-rate subscription: not spend |

_Source: Pi session JSONL usage records and Drill agent-invocation records, covering the task lifetime through report generation. Provider-recorded costs and list-price estimates are labeled separately; subscription usage is not represented as spend._

Pipeline

Updates from git push drill

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed ✅
  • 🚨 bin/sq-cost.sh:149 - The RETURN trap trap &#39;rm -f &#34;$enriched&#34;&#39; RETURN leaks out of cmd_report into its caller. cmd_report is invoked directly at top level for report (line 391) and via a subshell for publish (line 262), so the trap fires once while enriched is still in scope there. But cmd_task calls cmd_report &#34;$task_id&#34; --json in the same shell (line 288), so the trap fires a second time on cmd_task's return, after the cmd_report local is out of scope; under set -u this aborts with bin/sq-cost.sh: line 288: enriched: unbound variable and exit 1, even though the JSON was already written. At base, --json returned before the trap was set, so this is a regression. Reproduction (any found=true task, Pi-only or Drill): bash bin/sq-cost.sh task &lt;id&gt; --json prints the report then exits 1; bash bin/sq-cost.sh report &lt;id&gt; --json exits 0. Impact: task &lt;id&gt; --json is the documented equivalent of report --json (script header) and is the exact call bin/sq-trajectory.sh makes (&#34;$COST&#34; task &#34;$id&#34; --json ... || cost_json=&#39;{}&#39;), so callers that check status silently lose the report. No existing test executes task &lt;id&gt; --json with found=true, so the suite does not catch it. Fix: make the cleanup self-clearing and unset-safe, e.g. trap &#39;rm -f &#34;${enriched:-}&#34;; trap - RETURN&#39; RETURN, and add an executable regression asserting task &lt;id&gt; --json exits 0.
  • ℹ️ bin/sq-cost.sh:111 - cmd_report now always reads $HOME/.drill/state.sqlite when the Drill file exists, so report/publish and anything that calls them transitively depend on the developer's real Drill state unless SQUAD_DRILL_STATE is set. tests/sq-cost.test.sh handles this with a global export, but tests/sq-trajectory.test.sh invokes sq-cost.sh task &lt;id&gt; --json without overriding it; if a real run exists on branch sq/alpha (or sq/beta/sq/taskNN), the report becomes found=true and the test's .tokens.input.value==null contract breaks. This is currently masked by the trap bug above (that non-zero exit forces cost_json=&#39;{}&#39;), so fixing the trap will expose the non-hermetic read. Recommend sandboxing SQUAD_DRILL_STATE in that suite (or in the test runner) alongside the fix.

🔧 Fix: fix(sq-cost): self-clear report trap and isolate test Drill state
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/sq-cost.test.sh (all focused behavior tests pass, including Drill attribution, non-operator pipeline models, combined totals, cents formatting, flat-rate not-spend, message-level provider attribution, and Drill read-failure visibility)
  • bin/sq-cost.sh report squad-plan-execute-dispatch-contract --json against the real ~/.drill/state.sqlite and compared with jq -S to the intent-provided expected output (exact semantic match)
  • jq &#39;[paths(type==&#34;number&#34;) as $p | select($p[-1]|test(&#34;cost&#34;))]&#39; over target and real report JSON to confirm no numeric money fields remain
  • Standalone CLI harness building a synthetic operator Pi session plus Drill SQLite fixture, then running sq-cost.sh report pipeline-task (JSON and Markdown) from both base commit 080f6cb and target de9badf
  • bash tests/sq-cost.test.sh against a git archive extraction of base commit 080f6cb to confirm the new regression fails on the previous omission (exit 1 on the cents assertion)
  • git status --porcelain to confirm no residual worktree changes after testing
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@90sRehem

90sRehem commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Coding agent usage on this pull request

Contributor Agent Sessions Total tokens Estimated cost
Squad task squad-cost-report-models-and-pipeline Pi 1 24.9 million $0.33836732

Token and model breakdown

Model Input Output Cache read Cache write Total tokens Estimated cost
gpt-6-luna 557.8 thousand 79.7 thousand 24.3 million 0 24.9 million provider-recorded: 0.33836732

Source: Pi session JSONL usage records, covering the task lifetime from 2026-10-09T13:29:46.582Z through report generation. Costs are provider-recorded where available, otherwise list-price estimates; subscription usage is not represented as spend.

@90sRehem
90sRehem merged commit 1757d14 into main Oct 9, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant