Repository navigation
Document agent.maxToolResultBytes and how the agent recovers when a conversation outgrows the model's context window - #720
DavidCockerill wants to merge 5 commits into
Conversation
…onversation outgrows the model's context window Refs HarperFast/harper#3068 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refs HarperFast/harper#3068 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…han a page Refs HarperFast/harper#3068 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ase-note heading overpromising Refs HarperFast/harper#3068 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request updates the documentation and release notes for version 5.4.0 to introduce the maxToolResultBytes configuration option, which caps and pages agent tool results to prevent conversations from exceeding the model's context window. The review feedback suggests improving the phrasing of how tool results are cut for better readability and adding a version badge to the set_agent_config section to properly document the API change.
| Every model request replays the session's whole conversation, so a single large tool result can fill the model's context window, and then every later request is rejected. Two rules keep a session usable: | ||
|
|
||
| - **Tool results are capped where they are stored.** A result larger than [`agent.maxToolResultBytes`](../configuration/options.md#agent) (default 64 KiB) is cut to that size before it is added to `messages`, ending with a note that gives its original size and tells the model how to ask for less. Only the cut form is kept, so `get_agent_session` shows what the model saw. The file-reading tools return pages sized to fit under the cap, measured as JSON: `read_file` returns up to `lineCount` whole lines from `startLine`, and while the file continues it gives `nextLine` and `nextOffset`, which the agent passes back as `startLine` and `offset` to read the next page without rescanning the file. So it can work through a log of any size one page at a time. A line longer than a page comes back in parts, continued the same way, with every part but the last flagged `lineTruncated`. | ||
| - **A rejected request is retried once.** When the provider rejects a request because it does not fit the model's context window (detected for the OpenAI, Anthropic and Bedrock backends), the agent cuts to 2 KiB every result over 2 KiB in the most recent group of tool results that has one, keeping its beginning plus a note saying it was cut, records the cut in `messages`, and sends the request again. If nothing is left to cut, or the retry is rejected as well, the run ends `error` with a `lastError` that starts `The conversation no longer fits the model's context window`. Shorten the prompt, or start a new session. With fallback models configured, a context-window rejection from a fallback that follows an unrelated failure of the first model is reported as that first failure, and is not retried. |
| ### `set_agent_config` | ||
|
|
||
| Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only. | ||
| Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxToolResultBytes`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only. A `maxToolResultBytes` that is not an integer from `1024` to `1048576` is rejected the same way. |
There was a problem hiding this comment.
Since set_agent_config is an existing API surface and we are introducing a new configuration key (maxToolResultBytes) to its accepted parameters, we should include a component directly under the heading to denote this behavior change.
References
- When documenting behavior changes to an existing API surface (e.g., adding new fields to a response), use . Ensure it is placed standalone under headings rather than mid-sentence.
🚀 Preview DeploymentYour preview deployment is ready! 🔗 Preview URL: https://preview.harper-documentation.harperfabric.com/pr-720 This preview will update automatically when you push new commits. |
⊙ Problem
The built-in agent no longer stores tool results in full, pages
read_file, and retries once when a conversation outgrows the model's context window. None of that is documented, including the newagent.maxToolResultBytessetting (Cap the agent's tool results and page its file reads, so one large read no longer ends the session (harper#3117), for agent: no cap on tool-result size — one large read kills the session (harper#3068)).Refs HarperFast/harper#3068
💡 Solution
reference/configuration/options.md:maxToolResultBytesin the agent example and option list (default, range, startup fallback, page size), and in the list of keysset_agent_configcan change.reference/operations-api/operations.md: a new section, When a conversation outgrows the model's context window. It covers the cap,read_file's cursor, and the single retry, with its limits: best effort, the fallback-candidate case, and recovering a session broken by an earlier version. Alsoset_agent_config's new key and its 400, and that a run keeps the cap it started with.release-notes/v5-lincoln/5.4.md: a Built-in Agent entry.✅ Verification
Two cross-model review rounds (Codex + Gemini + Harper adjudication) checked every claim against the code PR. They caught overclaims about paging and recovery, which are corrected here, and a page-budget gap that is fixed in the code PR.
prettier --checkpasses on the changed files.🤖 Generated by Claude (Anthropic); posted via @DavidCockerill.
Related PRs: #671 independent, #676 independent, #683 independent, #691 independent, #697 independent, #709 overlaps (5.4 release notes; rebase conflict only), #712 overlaps (same agent option list, set_agent_config key lists and 5.4 Built-in Agent heading), #713 independent, #710 independent, #707 independent
Complexity: easy
Review-Coverage: authored=claude; ran=gemini,codex; adjudicated=domain; blocked=cursor-composer(not-installed); declined=cursor-grok,cursor-kimi,cursor-muse; rounds=2; full=1 @ 5207085
Review-Attention: skim ~2m (decisions: retry-scope, fix-layer-for-page-fit) @ 5207085