Skip to content

Fix native system prompts and expose cache usage - #73

Merged
clifton merged 2 commits into
mainfrom
agent/system-prompts-and-caching
Jul 30, 2026
Merged

clifton merged 2 commits into
mainfrom
agent/system-prompts-and-caching

Conversation

@clifton

@clifton clifton commented Jul 30, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes #72.

Request::with_system previously concatenated the system text into the user
prompt for materialize, generate, streaming, and run without tools. Only
the tool loop used a native system field/message. Besides changing provider
instruction semantics, that made the supposedly stable prefix vary with every
user prompt and prevented compatible prompt caches from matching it.

This PR:

  • routes system instructions through every built-in provider's native request
    shape for every request-builder terminal;
  • preserves the exact [system, user] prefix across structured-output retries;
  • keeps AnyClient, media requests, streaming, attempt ledgers, and tool-less
    run on the same path;
  • retains the old concatenation behavior as the default hook implementation for
    third-party LLMClient implementations, so existing custom clients continue
    to compile and behave as before;
  • exposes provider-reported prompt-cache reads and writes through
    TokenUsage::{cached_input_tokens, cache_write_input_tokens} and aggregates
    them in RunUsage;
  • fixes Anthropic accounting when cache metadata is present: its base
    input_tokens excludes cache reads and creations, so all three counters are
    now combined into rstructor's total input count;
  • documents cache behavior and deliberately does not enable paid/explicit cache
    creation by default.

Provider request shapes

Provider System instruction
OpenAI and OpenAI-compatible endpoints Initial messages[].role = "system"
xAI/Grok Initial messages[].role = "system"
Anthropic Top-level system
Gemini Top-level systemInstruction

Usage example

use rstructor::{LLMClient, OpenAIClient, RequestExt};

const RISK_POLICY: &str =
    "Use the fund's base currency. Report exposure as a multiple of NAV.";

let client = OpenAIClient::from_env()?;
let report = client
    .with_system(RISK_POLICY)
    .materialize_with_attempts::<Portfolio>(daily_positions)
    .await?;

if let Some(usage) = report.cumulative_usage {
    println!(
        "{} of {} input tokens came from cache",
        usage.cached_input_tokens,
        usage.input_tokens,
    );
}

The public request-builder syntax is unchanged. The wire behavior changes from
one combined user message to a native system instruction plus a user message.

Prompt-caching scope

OpenAI, Gemini, and xAI can use implicit prefix caching for eligible requests,
so preserving a stable native system prefix is immediately useful. Anthropic
requires cache-control configuration to create caches. This PR does not
silently enable Anthropic cache writes, OpenAI cache-routing keys/options, or
Gemini explicit cache objects because those controls can affect billing,
retention, and routing behavior.

Cache read/write counters are subsets of input_tokens; total_tokens() does
not double-count them.

Provider references:

Regression coverage

  • 15 mocked request-shape tests cover all four providers, structured and raw
    terminals, streaming, media, AnyClient, attempt ledgers, and run without
    tools.
  • Retry-history unit coverage verifies that the system/user prefix remains byte
    stable while assistant output and validation feedback are appended.
  • Four provider-shaped response fixtures verify OpenAI, Anthropic, Gemini, and
    xAI cache-token accounting.
  • Run-level tests cover per-model cache aggregation and saturating overflow.

Validation

  • cargo fmt --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo test --lib --all-features (299 passed)
  • cargo test --test system_prompt_request_tests --all-features (15 passed)
  • cargo test --test prompt_cache_usage_tests --all-features (4 passed)
  • cargo test --test documentation_gallery_tests --all-features (5 passed)
  • cargo test --doc --all-features (82 passed)
  • schema-only builds with derive and derive,mock

@djmaze

djmaze commented Jul 30, 2026

Copy link
Copy Markdown

Thanks, this seems to work for me (with an OpenAI-compatible provider).

@clifton
clifton merged commit fe36a34 into main Jul 30, 2026
9 checks passed
@clifton
clifton deleted the agent/system-prompts-and-caching branch July 30, 2026 11:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Any reason why generate and and materialize do not write real system prompts?

2 participants