Skip to content

fix(tools): move code execution behind the built-in tool seam - #1983

Merged
daavoo merged 2 commits into
mainfrom
refactor/code-execution-native-rendering
Oct 6, 2026
Merged

daavoo merged 2 commits into
mainfrom
refactor/code-execution-native-rendering

Conversation

@daavoo

@daavoo daavoo commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Description

When Otari runs a tool on a caller's behalf and the caller asked for it in a provider's own words, Otari answers in that provider's vocabulary: Anthropic's server_tool_use blocks on Messages, OpenAI's output items on Responses. Each built-in tool says how it does that through one registry, except code execution, which the request loops still special-cased by name. This puts code execution behind the same registry as web search, so the loops hold no tool-specific branch.

One thing callers can observe changes, deliberately. On Responses, when one model turn runs several tools, each code_interpreter_call item now sits where its call ran, among the other native items such as web_search_call. Before, every code_interpreter_call in the turn came after all of that turn's search items, so a turn that searched, ran code, searched and ran code again was reported as search, search, code, code. Now it is reported as search, code, search, code, in both the streaming and non-streaming responses. Nothing else changes: the same items are emitted, with the same content, to the same callers.

To get that ordering, the non-streaming Responses loop now records each call's native items right after the call runs, which is what the Messages loop and the Responses stream already did.

This is the code-execution step of #1927.

How to test it locally

  • tests/unit/test_mcp_loop_responses.py::test_native_items_keep_the_order_their_calls_ran_in and ::test_stream_native_items_keep_the_order_their_calls_ran_in pin the ordering above. Both fail on main (the search items come first) and pass here.
  • tests/unit/test_builtin_tools.py::test_code_execution_is_rendered_only_in_the_dialect_whose_keyword_declared_it pins which declaration asks for which rendering: Anthropic's dated code_execution_<date> on Messages, OpenAI's code_interpreter on Responses, and neither for otari_code_execution or the bare code_execution.
  • Every test file the issue lists, and every other one touching native items, code execution or web search, passes locally: 3748 passed.
  • uv run --frozen --no-dev python scripts/oss_edition_smoke.py and scripts/hybrid_edition_smoke.py pass, as do the architecture check, ruff check, ruff format --check and mypy.

The wire contract to the execution backend is untouched, and _pipeline.py changes only in ToolContext.native_tools. #1616 touches the same loop modules, so whichever lands second rebases.

PR Type

  • New Feature
  • Bug Fix
  • Refactor
  • Documentation
  • Infrastructure / CI

Relevant issues

Fixes #1197

Part of #1927.

Checklist

  • I understand the code I am submitting.
  • I have added or updated tests that cover my change (tests/unit, tests/integration).
  • I ran the Definition of Done checks locally (make lint, make typecheck, make test). The Python checks and every affected test file; CI runs the full suite.
  • Documentation was updated where necessary. docs/tools.md says where a code_interpreter_call item lands.
  • If the API contract changed, I regenerated the OpenAPI spec (uv run python scripts/generate_openapi.py). Not applicable: no route or schema changed.
  • If this changes a rule in ARCHITECTURE.md or scripts/check_architecture.py, the description names the rule and says why. Not applicable.

AI Usage

  • No AI was used.
  • AI was used for drafting/refactoring.
  • This is fully AI-generated.

AI Model/Tool used:
Claude Code (Claude Opus 5.5)

Any additional AI details you'd like to share:
Implemented, tested and opened by an agent at the author's direction; the author reviewed the approach.

NOTE:
When responding to reviewer questions, please respond yourself rather than copy/pasting reviewer comments into an AI and pasting back its answer. We want to discuss with you, not your AI :)

  • I am an AI Agent filling out this form (check box if true)

Summary

  • Added native code-execution renderings for the Messages and Responses APIs. The tool loops now use the shared rendering registry instead of handling code execution separately.
  • Removed code-execution-specific dispatch from both loops. Moved the code-interpreter call ID prefix to the tools package.
  • Documented that Responses code-interpreter items appear in call order alongside other native items.

Technical notes

  • Responses records each code_interpreter_call item when its call runs, preserving order among native items.
  • The change preserves the execution-backend wire contract.
  • Added tests for dialect-specific declarations and call ordering in streamed and non-streamed Responses, plus interleaved Messages tool-use blocks.

Code execution's registry entry now declares a native rendering for
Messages and Responses, in services/tools/_code_execution_messages.py and
_code_execution_responses.py, like web search's. The dialect loops lose
their code-execution branches, and ToolContext.native_tools resolves code
execution through native_rendering(...).declared(...) like every other
built-in tool.

The non-streaming Responses loop now collects native items right after
each call, as the Messages loop and the Responses stream already did. A
code_interpreter_call item therefore sits at its own call's place among
the batch's native items, in the order the calls ran, where before every
one of them came after the batch's web_search_call items.

Fixes #1197
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:16 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:16 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:16 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:16 — with GitHub Actions Active
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: mozilla-ai/otari/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: b26e5657-6402-467f-9461-421b66f7976b
📥 Commits

Reviewing files that changed from the base of the PR and between cc01de3 and 9626c34.

📒 Files selected for processing (6)
  • docs/tools.md
  • src/gateway/services/sandbox_backend.py
  • src/gateway/services/tools/_code_execution_declarations.py
  • src/gateway/services/tools/_code_execution_messages.py
  • src/gateway/services/tools/_code_execution_responses.py
  • tests/unit/test_mcp_loop_messages.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • docs/tools.md

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review.


Walkthrough

Code execution now uses dialect-specific native renderers for Messages and Responses. The Responses loop collects rendered items during tool execution, preserving their order with other native items in non-streaming output and streamed events.

Changes

Native code execution rendering

Layer / File(s) Summary
Register dialect-specific renderers
src/gateway/services/tools/_code_execution_*.py, src/gateway/services/tools/_code_execution_tool.py, src/gateway/services/tools/__init__.py, src/gateway/api/routes/_pipeline.py, tests/unit/test_builtin_tools.py
Code execution now has Messages and Responses renderers. The tool registry supplies each renderer for its dialect. ToolContext.native_tools selects declared tools through the rendering interface, and tests cover dialect selection.
Render Messages code execution
src/gateway/services/mcp_loop_messages.py, tests/unit/test_mcp_loop_messages.py, tests/unit/test_messages_minted_block_stripping.py
The Messages loop delegates native block generation to the Messages renderer. The renderer emits server-tool-use and result blocks for drained executions. Tests cover block order for interleaved search and code-execution calls.
Collect Responses items in call order
src/gateway/services/mcp_loop_responses.py, src/gateway/api/routes/responses.py, src/gateway/services/sandbox_backend.py, docs/tools.md, tests/unit/test_mcp_loop_responses.py, tests/unit/test_responses_produced_images.py
The Responses loop collects native renderings during each call rather than reconstructing them afterward. Tests cover ordering for interleaved web-search and code-execution calls in regular output and streamed events. The documentation states that one code_interpreter_call item is emitted per run and ordered with other native items.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Suggested reviewers: peteski22

Merge Risk: ⚪ Minimal · up to 9626c

Responses now preserves code-interpreter items in call order alongside other native items, while Messages uses its dialect-specific renderer. No material merge risk was identified.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 60.87% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 46 functions across 15 files. (1 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed #1197 is met. The code-execution tool declares Messages and Responses renderings, and the tests cover the provider-specific declaration keywords. The loop changes use the rendering seam and record nat…
Out of Scope Changes check ✅ Passed The changes support #1197. The documentation, renderer export, declaration type update, buffer-consumption comment, import updates, and added tests support the rendering seam or its call-order behavio…
Title check ✅ Passed The title uses the fix Conventional Commit prefix with a scope, uses imperative wording, stays under 70 characters, and describes the code-execution refactor.
Description check ✅ Passed The description clearly explains the change and its user-visible effect, provides local test steps and reported results, identifies the PR type and related issues, and completes the checklist and AI u…
Full details: Docstring Coverage

Explanation

Docstring coverage is 60.87% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 46 functions across 15 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
✨ Simplify code
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@reviewsaur

reviewsaur Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

🦕 Reviewsaur Quiz

A review comprehension quiz has been generated for this PR.

Take the quiz

Attempt 1 | (0/1 approval) | This link is for this PR's reviewers and expires when the PR is closed.

Tip: To require this quiz before merging, enable it as a required status check.

PR Walkthrough — what this change does, how it works, and what it influences

Complexity: normal — Diff touches the tool registry seam, both API loops, pipeline native_tools logic, sandbox docstring, docs and multiple test files.

What this change does

Moves code execution's native renderings for Dialect.MESSAGES and Dialect.RESPONSES into the BuiltinTool registry via new _code_execution_messages.py and _code_execution_responses.py modules, exactly as web search does. Removes the name-based branches (_native_code_execution_blocks, _code_interpreter_items, native_code_execution_dialect special case) from mcp_loop_messages.py, mcp_loop_responses.py and _pipeline.py.native_tools. Non-streaming Responses paths now pass native_items into _execute_function_calls so each call's items are appended immediately after it runs.

How it works

  1. BuiltinTool registers the two new RENDERING objects under Dialect keys
  2. native_rendering(name, dialect) returns the rendering; declared(entry) decides inclusion in native_tools
  3. rendering.ran(call, pool) drains take_executions() and produces the provider-shaped blocks/items
  4. Messages and Responses loops replace their inline code-execution handling with rendering.ran/refused calls
  5. Responses non-stream path records native_items right after each owned call instead of after the whole batch
  6. streaming path and Messages path already did per-call recording, so order among mixed web_search + code_execution calls is now uniform

Where it sits

Entry points: responses_tool_loop / responses_tool_loop_stream · anthropic_tool_loop · ToolContext.native_tools
Changed here: gateway/services/tools/_code_execution_tool.py · gateway/services/tools/_code_execution_messages.py · gateway/services/tools/_code_execution_responses.py · gateway/services/tools/init.py · gateway/services/mcp_loop_messages.py · gateway/services/mcp_loop_responses.py · gateway/api/routes/_pipeline.py
Feeds into: Messages API responses containing server_tool_use blocks · Responses API output containing code_interpreter_call items · sandbox_backend.take_executions callers

What this influences

  • web_search native rendering path
  • Dialect enum and native_rendering lookup
  • ToolUseBudget refusal paths
  • docs/tools.md description of item ordering
  • test_mcp_loop_responses ordering tests
  • test_builtin_tools dialect declaration test

File by file

  • src/gateway/services/tools/_code_execution_tool.py — Adds native={MESSAGES: ..., RESPONSES: ...} to the BuiltinTool definition.
  • src/gateway/services/tools/_code_execution_messages.py — New file containing MessagesCodeExecutionRendering with declared/ran/refused using the Anthropic block shapes.
  • src/gateway/services/tools/_code_execution_responses.py — New file containing ResponsesCodeExecutionRendering and CODE_INTERPRETER_CALL_ID_PREFIX with the code_interpreter_call item builder.
  • src/gateway/services/tools/__init__.py — Exports CODE_INTERPRETER_CALL_ID_PREFIX from the new responses module.
  • src/gateway/api/routes/_pipeline.py — native_tools now uses only native_rendering(...).declared; removes native_code_execution_dialect property and import.
  • src/gateway/services/mcp_loop_messages.py — Deletes _native_code_execution_blocks and the name check in _native_blocks_for_call; now routes everything through native_rendering.
  • src/gateway/services/mcp_loop_responses.py — Deletes code_interpreter* helpers and code_interpreter_items buffer; _execute_function_calls and stream path now receive/append native_items per call.
  • src/gateway/services/sandbox_backend.py — Updates take_executions docstring to state it must be drained after each call.
  • docs/tools.md — Updates Responses paragraph to say code_interpreter_call items appear in call order among other native items.

Generated from this PR's diff — it describes the change and its immediate connections, not the full repository.

… the seam

Document on take_executions that a loop must read it after each call,
because a code-execution rendering takes the whole buffer as that call's.
Add a Messages test interleaving searches and code runs that holds the
same ordering the Responses tests do. Accept a Mapping in
native_code_execution_dialect so the renderings stop copying the entry,
and rewrap the docs paragraph.
@daavoo daavoo changed the title refactor(tools): move code execution behind the built-in tool seam fix(tools): move code execution behind the built-in tool seam Oct 6, 2026
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:52 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:52 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:52 — with GitHub Actions Active
@daavoo
daavoo deployed to integration-tests October 6, 2026 10:52 — with GitHub Actions Active
@reviewsaur

reviewsaur Bot commented Oct 6, 2026

Copy link
Copy Markdown

🦕 Quiz Passed!

@daavoo scored 75% on attempt 1. (1/1 approval)

All required reviewers have passed. The review check has been marked as successful.

celebratory dinosaur

@daavoo
daavoo merged commit 2c120d7 into main Oct 6, 2026
27 checks passed
@daavoo
daavoo deleted the refactor/code-execution-native-rendering branch October 6, 2026 11:17

This branch was successfully deployed

1 active deployment
integration-tests — 9626c340 Deployed Oct 6, 2026 by daavoo via test-integration (1/4) #3288
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Move sandbox (code execution) behind the seam

1 participant