Fix native system prompts and expose cache usage - #73
Merged
Merged
Conversation
|
Thanks, this seems to work for me (with an OpenAI-compatible provider). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #72.
Request::with_systempreviously concatenated the system text into the userprompt for
materialize,generate, streaming, andrunwithout tools. Onlythe tool loop used a native system field/message. Besides changing provider
instruction semantics, that made the supposedly stable prefix vary with every
user prompt and prevented compatible prompt caches from matching it.
This PR:
shape for every request-builder terminal;
[system, user]prefix across structured-output retries;AnyClient, media requests, streaming, attempt ledgers, and tool-lessrunon the same path;third-party
LLMClientimplementations, so existing custom clients continueto compile and behave as before;
TokenUsage::{cached_input_tokens, cache_write_input_tokens}and aggregatesthem in
RunUsage;input_tokensexcludes cache reads and creations, so all three counters arenow combined into rstructor's total input count;
creation by default.
Provider request shapes
messages[].role = "system"messages[].role = "system"systemsystemInstructionUsage example
The public request-builder syntax is unchanged. The wire behavior changes from
one combined user message to a native system instruction plus a user message.
Prompt-caching scope
OpenAI, Gemini, and xAI can use implicit prefix caching for eligible requests,
so preserving a stable native system prefix is immediately useful. Anthropic
requires cache-control configuration to create caches. This PR does not
silently enable Anthropic cache writes, OpenAI cache-routing keys/options, or
Gemini explicit cache objects because those controls can affect billing,
retention, and routing behavior.
Cache read/write counters are subsets of
input_tokens;total_tokens()doesnot double-count them.
Provider references:
Regression coverage
terminals, streaming, media,
AnyClient, attempt ledgers, andrunwithouttools.
stable while assistant output and validation feedback are appended.
xAI cache-token accounting.
Validation
cargo fmt --checkcargo clippy --all-targets --all-features -- -D warningscargo test --lib --all-features(299 passed)cargo test --test system_prompt_request_tests --all-features(15 passed)cargo test --test prompt_cache_usage_tests --all-features(4 passed)cargo test --test documentation_gallery_tests --all-features(5 passed)cargo test --doc --all-features(82 passed)deriveandderive,mock