Skip to content

Support Codex CLI image attachments with vLLM #251

Description

@franciscojavierarceo

Outcome

Codex CLI users can attach local images and use view_image through Agentic API with a supported vision-capable vLLM model. Images must reach inference and survive Responses continuation without being silently removed or treated as base64 text for context budgeting.

Related: #54 (broader Codex integration). This tracker scopes attachment handling only.

Findings

Source investigation used OpenAI Codex main at 6af345407d9c2a568da9d01b6c4b81a9e61495c0 and Agentic API at c04ac196f44133c37be652c4ee3f777db32ae6df.

  • Codex converts local image attachments to inline data URLs in input_image; view_image returns structured image content in a function call output. Ordinary file mentions identify local paths for client-executed tools.
  • The gateway already models images in messages and tool call outputs, but the launcher hardcodes text-only model metadata and the HTTP catalog relies on upstream capability metadata that may be absent.
  • HTTP and WebSocket input limits are fixed at 10 MiB. Automatic compaction estimates tokens from serialized JSON, including base64 payloads.
  • User-message input_file falls through to Unknown; this is a separate validation gap, not a requirement for the ordinary CLI file-mention workflow.

Scope and sequencing

Initial CLI image milestone: resolve capabilities consistently across catalog and launcher paths, then prove attachment and view_image behavior with deterministic coverage and a pinned live vLLM/model check. Existing request limits remain applicable until the limits sub-issue lands.

Hardening: configurable transport limits and media-aware compaction estimation. Compaction support is not complete until that sub-issue lands.

Related API correctness: explicitly reject unsupported file content on typed execution paths without changing transparent proxy semantics. This can ship independently and does not block the initial CLI image milestone.

Sub-issues will define implementation boundaries and acceptance criteria.

Explicitly out of scope

  • A Codex fork or changes to Codex attachment serialization.
  • Files API endpoints, multipart uploads, blob storage, file-ID resolution, PDF/DOCX ingestion, OCR, or document rendering services.
  • Audio/video, image generation, original-detail guarantees, or a generic model-capability platform.
  • Assuming vision capability from a model name or claiming all vLLM model/template combinations work.

Completion

All linked sub-issues are complete; docs identify the tested Codex version, vLLM version, model/template, image limits, and unsupported file behavior. A pinned live test must establish image understanding; replay tests alone establish transport behavior, not model capability.

Work breakdown

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions