Skip to content

Document agent.maxToolResultBytes and how the agent recovers when a conversation outgrows the model's context window - #720

Draft
DavidCockerill wants to merge 5 commits into
mainfrom
david/agent-tool-result-cap
Draft

DavidCockerill wants to merge 5 commits into
mainfrom
david/agent-tool-result-cap

Conversation

@DavidCockerill

@DavidCockerill DavidCockerill commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

⊙ Problem

The built-in agent no longer stores tool results in full, pages read_file, and retries once when a conversation outgrows the model's context window. None of that is documented, including the new agent.maxToolResultBytes setting (Cap the agent's tool results and page its file reads, so one large read no longer ends the session (harper#3117), for agent: no cap on tool-result size — one large read kills the session (harper#3068)).

❓ Your call: badged v5.4.0 to match Document agent.maxTokens, and that a cut-off or empty agent reply ends the run error (#712); it follows the code PR's milestone if that changes.

Refs HarperFast/harper#3068

💡 Solution

❓ Your call: whichever of this PR and #712 merges second keeps both new keys in the three lists and a single ## Built-in Agent heading in the 5.4 notes.

✅ Verification

Two cross-model review rounds (Codex + Gemini + Harper adjudication) checked every claim against the code PR. They caught overclaims about paging and recovery, which are corrected here, and a page-budget gap that is fixed in the code PR. prettier --check passes on the changed files.

🤖 Generated by Claude (Anthropic); posted via @DavidCockerill.

Related PRs: #671 independent, #676 independent, #683 independent, #691 independent, #697 independent, #709 overlaps (5.4 release notes; rebase conflict only), #712 overlaps (same agent option list, set_agent_config key lists and 5.4 Built-in Agent heading), #713 independent, #710 independent, #707 independent
Complexity: easy

Review-Coverage: authored=claude; ran=gemini,codex; adjudicated=domain; blocked=cursor-composer(not-installed); declined=cursor-grok,cursor-kimi,cursor-muse; rounds=2; full=1 @ 5207085

Review-Attention: skim ~2m (decisions: retry-scope, fix-layer-for-page-fit) @ 5207085

DavidCockerill and others added 5 commits October 6, 2026 16:42
…onversation outgrows the model's context window

Refs HarperFast/harper#3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refs HarperFast/harper#3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…han a page

Refs HarperFast/harper#3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ase-note heading overpromising

Refs HarperFast/harper#3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the documentation and release notes for version 5.4.0 to introduce the maxToolResultBytes configuration option, which caps and pages agent tool results to prevent conversations from exceeding the model's context window. The review feedback suggests improving the phrasing of how tool results are cut for better readability and adding a version badge to the set_agent_config section to properly document the API change.

Every model request replays the session's whole conversation, so a single large tool result can fill the model's context window, and then every later request is rejected. Two rules keep a session usable:

- **Tool results are capped where they are stored.** A result larger than [`agent.maxToolResultBytes`](../configuration/options.md#agent) (default 64 KiB) is cut to that size before it is added to `messages`, ending with a note that gives its original size and tells the model how to ask for less. Only the cut form is kept, so `get_agent_session` shows what the model saw. The file-reading tools return pages sized to fit under the cap, measured as JSON: `read_file` returns up to `lineCount` whole lines from `startLine`, and while the file continues it gives `nextLine` and `nextOffset`, which the agent passes back as `startLine` and `offset` to read the next page without rescanning the file. So it can work through a log of any size one page at a time. A line longer than a page comes back in parts, continued the same way, with every part but the last flagged `lineTruncated`.
- **A rejected request is retried once.** When the provider rejects a request because it does not fit the model's context window (detected for the OpenAI, Anthropic and Bedrock backends), the agent cuts to 2 KiB every result over 2 KiB in the most recent group of tool results that has one, keeping its beginning plus a note saying it was cut, records the cut in `messages`, and sends the request again. If nothing is left to cut, or the retry is rejected as well, the run ends `error` with a `lastError` that starts `The conversation no longer fits the model's context window`. Shorten the prompt, or start a new session. With fallback models configured, a context-window rejection from a fallback that follows an unrelated failure of the first model is reported as that first failure, and is not retried.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The phrasing "cuts to 2 KiB every result over 2 KiB" is slightly awkward. Rephrasing it to "cuts every result over 2 KiB to 2 KiB" improves readability and flow.

### `set_agent_config`

Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only.
Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxToolResultBytes`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only. A `maxToolResultBytes` that is not an integer from `1024` to `1048576` is rejected the same way.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since set_agent_config is an existing API surface and we are introducing a new configuration key (maxToolResultBytes) to its accepted parameters, we should include a component directly under the heading to denote this behavior change.

References
  1. When documenting behavior changes to an existing API surface (e.g., adding new fields to a response), use . Ensure it is placed standalone under headings rather than mid-sentence.

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

🚀 Preview Deployment

Your preview deployment is ready!

🔗 Preview URL: https://preview.harper-documentation.harperfabric.com/pr-720

This preview will update automatically when you push new commits.

This branch was successfully deployed

1 active deployment
pr-720 — 52070856 Deployed Oct 8, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant