Skip to content

feat: add configurable execution token budgets - #2852

Open
qinxuye wants to merge 1 commit into
xorbitsai:mainfrom
qinxuye:feat/execution-token-budget
Open

qinxuye wants to merge 1 commit into
xorbitsai:mainfrom
qinxuye:feat/execution-token-budget

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add an execution-scoped token budget shared by nested agents, with checkpoint-aware pause/resume accounting and guards before model and tool calls.
  • Add configurable soft-warning percentages and deterministic hard-stop handoff: preserve already-created, registered files and report partial or blocked completion without spending another model call.
  • Add administrator defaults and maximums, personal overrides, bilingual settings controls, and a typed policy-resolver extension. Wire the policy into normal tasks, preview, builder chat, and workforce creation.

Behavior and scope

  • No token cap is enabled by default. With no effective cap, there is no budget warning or stop, and no budget-only checkpoint write.
  • Blank personal settings inherit administrator/environment settings; administrator maximums still apply. The soft-warning percentage defaults to 80% and is configurable.
  • Accounting uses provider-reported input plus output tokens, including cached input once. Existing billing and quota behavior remains separate.
  • Enforcement happens at call boundaries. In-flight calls can exceed the configured cap; a final answer already produced by a paid call is retained. This is not an exact billing cap or a guarantee of lower token usage.
  • Budget termination preserves available registered output links, identifies incomplete work, and does not claim that existing files have passed content validation.
  • No documentation files or experimental harness scripts are included.

Validation

  • Backend regression suite: 2,932 passed across tests/core/agent, configuration, token accounting, execution-budget web policy, workforce creation, and service-manager initialization tests.
  • Frontend regression suite: 49 passed across settings, execution budgets, user preferences, and translations.
  • Pre-commit checks passed, including Python lint/format/type checks and frontend lint/TypeScript checks.
  • Added regression coverage ensuring unlimited executions do not create an extra budget-only tail checkpoint.

Real-model task checks during development

Ran real saved tasks with the same synthetic business fixture to generate an XLSX workbook and CSV. The enabled runs deliberately used a 25,000-token cap and 30% soft threshold to exercise the mechanism; these are test values, not proposed defaults.

Model Budget Model calls Reported tokens Outcome
GPT-5.6 Luna Disabled 2 20,602 XLSX and CSV delivered
GPT-5.6 Luna 25,000 / 30% 3 31,496 XLSX and CSV delivered; final in-flight call exceeded the cap
DeepSeek Flash Disabled 6 65,178 XLSX and CSV delivered
DeepSeek Flash 25,000 / 30% 3 35,316 Stopped with XLSX available and CSV incomplete; reported partial completion

Enabled runs emitted one warning each; disabled controls emitted none. Verified accounting, call-boundary stopping, downloadable artifacts, CSV values, and XLSX source rows/formulas. Native Office recalculation and visual rendering were not part of these checks. An initial run with a mismatched Python environment was excluded and rerun with the task environment aligned.

These are single-run functional checks, not statistical performance results. They demonstrate enforcement and partial delivery, not a reliable token-saving effect from the soft warning.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants