fix(tools): move code execution behind the built-in tool seam - #1983
Conversation
Code execution's registry entry now declares a native rendering for Messages and Responses, in services/tools/_code_execution_messages.py and _code_execution_responses.py, like web search's. The dialect loops lose their code-execution branches, and ToolContext.native_tools resolves code execution through native_rendering(...).declared(...) like every other built-in tool. The non-streaming Responses loop now collects native items right after each call, as the Messages loop and the Responses stream already did. A code_interpreter_call item therefore sits at its own call's place among the batch's native items, in the order the calls ran, where before every one of them came after the batch's web_search_call items. Fixes #1197
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (6)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review. WalkthroughCode execution now uses dialect-specific native renderers for Messages and Responses. The Responses loop collects rendered items during tool execution, preserving their order with other native items in non-streaming output and streamed events. ChangesNative code execution rendering
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to Responses now preserves code-interpreter items in call order alongside other native items, while Messages uses its dialect-specific renderer. No material merge risk was identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 60.87% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 46 functions across 15 files. (1 skipped: 1 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
✨ Simplify code
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
🦕 Reviewsaur QuizA review comprehension quiz has been generated for this PR. Attempt 1 | (0/1 approval) | This link is for this PR's reviewers and expires when the PR is closed. Tip: To require this quiz before merging, enable it as a required status check. PR Walkthrough — what this change does, how it works, and what it influencesComplexity: normal — Diff touches the tool registry seam, both API loops, pipeline native_tools logic, sandbox docstring, docs and multiple test files. What this change doesMoves code execution's native renderings for Dialect.MESSAGES and Dialect.RESPONSES into the BuiltinTool registry via new _code_execution_messages.py and _code_execution_responses.py modules, exactly as web search does. Removes the name-based branches (_native_code_execution_blocks, _code_interpreter_items, native_code_execution_dialect special case) from mcp_loop_messages.py, mcp_loop_responses.py and _pipeline.py.native_tools. Non-streaming Responses paths now pass native_items into _execute_function_calls so each call's items are appended immediately after it runs. How it works
Where it sitsEntry points: responses_tool_loop / responses_tool_loop_stream · anthropic_tool_loop · ToolContext.native_tools What this influences
File by file
Generated from this PR's diff — it describes the change and its immediate connections, not the full repository. |
… the seam Document on take_executions that a loop must read it after each call, because a code-execution rendering takes the whole buffer as that call's. Add a Messages test interleaving searches and code runs that holds the same ordering the Responses tests do. Accept a Mapping in native_code_execution_dialect so the renderings stop copying the entry, and rewrap the docs paragraph.
🦕 Quiz Passed!@daavoo scored 75% on attempt 1. (1/1 approval) All required reviewers have passed. The review check has been marked as successful. |

Description
When Otari runs a tool on a caller's behalf and the caller asked for it in a provider's own words, Otari answers in that provider's vocabulary: Anthropic's
server_tool_useblocks on Messages, OpenAI's output items on Responses. Each built-in tool says how it does that through one registry, except code execution, which the request loops still special-cased by name. This puts code execution behind the same registry as web search, so the loops hold no tool-specific branch.One thing callers can observe changes, deliberately. On Responses, when one model turn runs several tools, each
code_interpreter_callitem now sits where its call ran, among the other native items such asweb_search_call. Before, everycode_interpreter_callin the turn came after all of that turn's search items, so a turn that searched, ran code, searched and ran code again was reported as search, search, code, code. Now it is reported as search, code, search, code, in both the streaming and non-streaming responses. Nothing else changes: the same items are emitted, with the same content, to the same callers.To get that ordering, the non-streaming Responses loop now records each call's native items right after the call runs, which is what the Messages loop and the Responses stream already did.
This is the code-execution step of #1927.
How to test it locally
tests/unit/test_mcp_loop_responses.py::test_native_items_keep_the_order_their_calls_ran_inand::test_stream_native_items_keep_the_order_their_calls_ran_inpin the ordering above. Both fail onmain(the search items come first) and pass here.tests/unit/test_builtin_tools.py::test_code_execution_is_rendered_only_in_the_dialect_whose_keyword_declared_itpins which declaration asks for which rendering: Anthropic's datedcode_execution_<date>on Messages, OpenAI'scode_interpreteron Responses, and neither forotari_code_executionor the barecode_execution.uv run --frozen --no-dev python scripts/oss_edition_smoke.pyandscripts/hybrid_edition_smoke.pypass, as do the architecture check,ruff check,ruff format --checkandmypy.The wire contract to the execution backend is untouched, and
_pipeline.pychanges only inToolContext.native_tools. #1616 touches the same loop modules, so whichever lands second rebases.PR Type
Relevant issues
Fixes #1197
Part of #1927.
Checklist
tests/unit,tests/integration).make lint,make typecheck,make test). The Python checks and every affected test file; CI runs the full suite.docs/tools.mdsays where acode_interpreter_callitem lands.uv run python scripts/generate_openapi.py). Not applicable: no route or schema changed.ARCHITECTURE.mdorscripts/check_architecture.py, the description names the rule and says why. Not applicable.AI Usage
AI Model/Tool used:
Claude Code (Claude Opus 5.5)
Any additional AI details you'd like to share:
Implemented, tested and opened by an agent at the author's direction; the author reviewed the approach.
NOTE:
When responding to reviewer questions, please respond yourself rather than copy/pasting reviewer comments into an AI and pasting back its answer. We want to discuss with you, not your AI :)
Summary
Technical notes
code_interpreter_callitem when its call runs, preserving order among native items.