Skip to content

Cap the agent's tool results and page its file reads, so one large read no longer ends the session - #3117

Draft
DavidCockerill wants to merge 9 commits into
mainfrom
david/agent-tool-result-cap
Draft

DavidCockerill wants to merge 9 commits into
mainfrom
david/agent-tool-result-cap

Conversation

@DavidCockerill

@DavidCockerill DavidCockerill commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

⊙ Problem

The built-in agent added every tool result to its transcript in full, and every model request replays the transcript. One large read_file of a log, the file an investigating agent most needs, pushed the next request past the model's context window. The provider rejected it, the run ended error, and every later prompt replayed the same oversized result, so the session could not be used again (agent: no cap on tool-result size — one large read kills the session (#3068)). Reproduced on 5230512f9: a 1 MB log read was stored as a 1,057,864-byte tool message, the next request was rejected, and a follow-up prompt on the same session was rejected the same way.

❓ Your call: the requirement, as harper#3068 specifies it, with two departures. The note after a cut result lives in the tool message itself, not a system message, because the Anthropic and Bedrock backends hoist inline system messages into the top-level system prompt, away from the result. The terminal reason is a distinct lastError, not a new structured field, matching how harper#3064 ends its runs. Milestoned v5.4 (the issue said v5.3), like its output-side twin harper#3064.

Closes #3068

Docs: Document agent.maxToolResultBytes and how the agent recovers when a conversation outgrows the model's context window (documentation#720)

💡 Solution

⚖️ Alternatives

  • Cap only, nothing else. Rejected: a cut log read gives the model no way to reach the rest of the file, and it cannot rescue a session that already holds an oversized result.
  • Cap when building the request, keep full results stored. Rejected: the session row is rewritten whole on every append, so stored multi-MB results multiply write cost, and the operator would see a transcript the model never saw.
  • Switch the agent to generate({ toolMode: 'auto' }). Rejected for now: the orchestrator has no destructive-call approval pause or per-turn persistence.
  • Line-only paging. Adopted after review: a byte cursor beside the line cursor. Without it every page rescanned the file, and a line longer than a page could not be read past its head, so a later write_file would drop its tail.

❓ Your call: the shrink rewrites the stored transcript. The cut results stay cut even if the retry still fails, and the originals survive only in the table's audit log. The alternative is to cut only the request, which then fails again on every later turn. Easy to reverse.

❓ Your call: the shrink takes the newest group of oversized results. That undoes the turn that crossed the limit and reaches a result stored before a later prompt. The cost: a session sitting at the limit loses each fresh result while older ones survive. Taking the oldest is a one-line change.

❓ Your call: the 65,536-byte default is lower than the issue's "a few hundred KB". It matches the orchestrator's default and keeps one result at about 16–20k tokens. The 1024 floor and 1 MiB ceiling stop a typo from switching the cap off.

❓ Your call: context-window detection reads provider wording (HTTP 400/413 plus a pattern, or OpenAI's context_length_exceeded code). If a provider changes its wording, recovery stops and the run ends error as before. A false positive costs one retry with the newest results cut.

❓ Your call: the system prompt now goes only in input.system. All four built-in backends honour it; a module backend that reads only messages would now lose it rather than receive it twice.

🔧 Changes

Product and architecture tour

What reaches the transcript

How big can one tool result get, and what does the model see when it is bigger?

Before After
A tool result was stored whole: a 1 MB log read became a 1,057,864-byte tool message, replayed on every later request. A result is cut to agent.maxToolResultBytes (65,536 by default) before it is stored, ending with a note naming its size and how to ask for less.

No observation over agent.maxToolResultBytes is appended to the transcript: tool results, tool errors and unknown-tool replies alike.
The cap applies when a result is stored, not when a request is built, because the session row is rewritten whole on every append.

Reading a file a page at a time

How does the agent read a log far larger than its cap?

Example: paging a 1 MB log

The model calls read_file with only a path and gets the first 286 lines, about 32 KB, with nextLine: 287 and nextOffset: 32384. It passes those back as startLine and offset for the next page, so the file is not rescanned. A minified bundle's single long line comes back in parts, continued the same way.

Outcome: the whole file is reachable, every page fits the cap, and nothing is skipped.

A page's JSON-escaped text never exceeds half the cap, less the room its path and fields need, so the loop never cuts a page and its cursor survives.
Paging keeps the fs tools' key-material refusal from #3098, including for a page that falls inside a key.

When the provider says the request is too long

What happens when a conversation outgrows the model's context window anyway?

One retry after a context-window rejection

The backend flags the error, so the loop never parses provider wording; it cuts the newest large results in the stored transcript and asks once more.

sequenceDiagram
    participant Loop
    participant Backend
    participant Session
    Loop->>Backend: generate with the transcript
    Backend-->>Loop: rejected, contextWindowExceeded
    Loop->>Session: cut newest results over 2 KiB
    Loop->>Backend: generate again
    Backend-->>Loop: answer, or a second rejection
    Note over Loop: second rejection ends the run error
Loading

A context-window rejection never ends a run before the newest oversized results are cut and the request retried once.
The retry is best effort: a long prompt or history still ends the run error, with a lastError that says to shorten the prompt or start a new session.

Every changed file:

✅ Verification

End-to-end route: live smoke with recorded evidence. Harper booted from dist/ with a fake OpenAI-compatible backend that returns OpenAI's real 400 context_length_exceeded envelope for any request over a byte limit, and a 1 MB logs/big.log the model reads with read_file.

Run Limit Result
base 5230512f9 300 KB tool message stored at 1,057,864 bytes; request 2 rejected; session error; a second prompt rejected the same way
this branch 300 KB read_file returns a 33 KB page with nextLine/nextOffset; both prompts completed; request 1 has one system message (24.3 KB, was 30.0 KB)
this branch 45 KB request 2 (58,024 bytes) rejected through the real OpenAIBackend; the page shrunk to 2,046 bytes; the retry (26,860 bytes) completed; the same on the second prompt
this branch, on the base run's dead session 300 KB the stored 1,057,864-byte result shrunk to 2,046 bytes on the first rejection; the session completed and paged the log

Unit tests (all pass on the final head):

Fails on base: the new tests, applied to 5230512f9, fail at their own assertions there. The loop tests die with the unrecovered prompt is too long, the system-prompt test sees ['system', 'user'], and the paging tests hit the old 5 MiB refusal.

After merging main (4e2148d0e, which brought #3098 into the same fs functions): unitTests/agent/*.test.js plus the backend, agentLoop and backendHelpers suites pass (500), and the 45 KB live smoke above repeats its result. A --final review of the merged state found that the look-back counted an END line just past the page start, so a page starting on a key's last body line was served; 1c3a31f33 counts only armor lines that begin before the start, with a test that fails without it. The full gates below ran before the merge; CI runs them on the merged head.

Gates before the merge (this Mac; every failure also fails on 5230512f9 in the same checkout, compared file by file):

  • test:unit:main, run on 5b71d2642 (the final head differs only in a test comment), with applicationSpawn.test.js excluded because it hangs (#2538): 6,703 passing, 32 failing. The failures are in globalIsolation, resolvePreload, RuntimeModuleTracker, cliOperations, processGroupReclaim, tokenAuthentication, jsLoader, OptionsWatcher, deployStagingRetention and deployActivation; base fails the same files (31 and 52 in two runs).
  • test:unit:resources, run on 474871b86 (later commits touch only agent/): 4,015 passing, 3 failing, all in sourceApplyConflictRetry, the same 3 on base.
  • test:integration:all, run on 474871b86: 2,270 passing, 21 failing, in isolated-application, shutdown-drain-e2e, log-rotation-write-path, rolling-restart, cert-key-reload and record-lock-concurrency. Base fails the identical 21 tests in those six files.
  • oxlint --deny-warnings, prettier --check and check:design-docs pass on the changed files.

❓ Your call: post-review deltas, not re-reviewed: grep_files's skippedFiles (about ten lines, for a finding from the pre-merge final round), and the armor look-back fix in 1c3a31f33 (five lines, the post-merge reviewer's own proposed fix), plus stubbing every Models sink in one test.
🤖 Generated by Claude (Anthropic); posted via @DavidCockerill.

Related PRs: #562 independent, #1770 overlaps (same backend throw sites; its retry wrapper must keep contextWindowExceeded and not retry a flagged 400), #1835 independent, #2155 independent, #2170 independent, #2436 independent, #2446 independent, #2532 overlaps (set_agent_config input schema would need maxToolResultBytes), #2649 independent, #2752 independent, #2763 independent, #2775 independent, #2901 independent, #2906 independent, #2930 independent, #2939 independent, #2986 independent, #2990 independent, #2994 independent, #3021 independent, #3030 independent, #3035 independent, #3057 independent, #3064 overlaps (same doRun, config and backend error lines; additive merge conflict), #3069 independent, #3065 independent, #3036 independent, #2981 independent, #3070 independent, #3072 independent, #3078 independent, #3079 independent, #3082 independent, #3083 independent, #3084 independent, #3086 independent, #3093 independent, #3107 independent, #3110 independent, #3113 overlaps (adjacent CONFIG_PARAMS lines in utility/hdbTerms.ts), #3098 overlaps (merged; same fs functions, conflict resolved in this branch), #3100 independent, #3103 independent, #3085 independent
Complexity: medium

Review-Coverage: authored=claude; ran=gemini,codex; adjudicated=domain; blocked=cursor-grok(not-installed); declined=cursor-composer,cursor-kimi,cursor-muse; rounds=4; full=3 @ 1c3a31f

Review-Attention: study ~9m (decisions: do-less-alternative, shrink-durability, read-file-unbounded, context-window-by-wording, absolute-paths-in-results; raised: open major) @ 1c3a31f

DavidCockerill and others added 9 commits October 6, 2026 16:37
…rom a context-window rejection

One large read_file/tail_file/http_fetch result was appended to the agent's
transcript verbatim, so the next model request overflowed the context window
and every later prompt replayed it: the session was dead.

- agent/loop.ts caps every observation at agent.maxToolResultBytes (default
  65536, 1024..1048576) with a marker saying how to ask for less, reusing the
  toolMode 'auto' serializer, now exact to the byte on UTF-8 boundaries.
- read_file pages by line (startLine/lineCount, nextLine/totalLines) and
  streams, so files over 5 MiB are readable; tail_file and grep_files return at
  most one page (half the cap) and say when they stopped early.
- The openai, anthropic and bedrock backends flag a context-window rejection on
  their error; the loop then shrinks the newest oversized tool results to 2 KiB
  and retries once, else ends the run with a lastError that says what to do.
- The system prompt is sent once (as `system`), not also as a message.
- A plain-object throw (Harper's BigInt toJSON) reaches the model as its message.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…indow rejection is recovered

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ges by their JSON-escaped length

Review fixes on the tool-result cap:

- read_file returns a byte cursor (offset/nextOffset) beside its line
  cursor. Passing both back resumes without rescanning the file, and a
  line longer than a page now comes back in parts instead of being
  skipped, so read-then-write_file no longer loses its tail.
- read_file, tail_file and grep_files measure a page JSON-escaped, so
  text dense in control characters or quotes no longer overflows the
  loop's cap and loses its cursor fields.
- tail_file uses bytesRead, so a file truncated between stat and read
  is not decoded from a zero-filled buffer.
- grep_files reports truncated when maxResults leaves matches unreturned,
  and does not split a surrogate pair when it cuts a line.
- An unknown-tool observation is capped like every other.
- A cancel that lands while the model is answering keeps the run aborted.
- A test drives the real OpenAI backend through Models into the loop.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o a long path cannot cut its cursor

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a cancelled agent run aborted

- grep_files at a small cap returned no result when its first match alone
  exceeded the page; it now returns that match cut to fit, with truncated.
- The loop rechecks the abort signal before recording completed and before
  each model request, so a cancel that lands during a session write keeps
  the run aborted.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t read as "nothing found"

grep_files never searches a file over 5 MiB, and a busy hdb.log rotates at
64 MB, so a grep of the logs could report no matches without having looked.
It now returns skippedFiles, and the description points at read_file, which
pages through any size.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
new Models(writer) replaced only the analytics writer; the metric emitter and
decision store still defaulted to Harper's real ones, so the test wrote into
the real analytics tables of the unit-test process.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves the agent filesystem tools against agent.configScope (#3098): paged
read_file and tail_file keep its scope rules and key-material refusal, and
both look back 32 KiB before their window for a BEGIN armor line, so a page
of key body alone, which carries no armor line, is refused too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
startsInsidePrivateKey reads 64 bytes past the page start so an armor line
cut there still matches, but it also counted an END line found in those 64
bytes, so a page starting within 39 bytes of a key's END line was returned:
up to 38 base64 characters of the key. Only armor lines that begin before the
page start now count.

Refs #3068

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a robust mechanism to prevent agent sessions from failing due to oversized tool results or context-window overflows. It adds a configurable maxToolResultBytes parameter (defaulting to 64 KiB) to cap the size of tool results appended to the transcript. The filesystem tools (read_file, grep_files, and tail_file) have been updated to dynamically page and size their outputs to fit within this budget. Additionally, the agent loop now detects context-window rejections from LLM providers (OpenAI, Anthropic, Bedrock), automatically shrinks the newest oversized tool results in the transcript, and retries the request once. Comprehensive unit tests have been added to validate these paging, capping, and recovery behaviors. No review comments were provided, so there is no feedback to address.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant