Environment
- Environment: Cloud Staging
- Model:
qwen/qwen3.7-max
- API: Chat Completions
- Affected requests: Streaming
Description
Two streaming requests to qwen/qwen3.7-max completed successfully, but no usage was recorded for either request.
Both requests:
- Returned HTTP
200
- Completed with
finish_reason: "stop"
- Received the final
[DONE] event
- Produced valid model responses
However:
- Neither request appeared in API Key Usage Details
- No corresponding usage / credit deduction was observed
Expected Result
Every successfully completed inference request should be recorded in usage metering and reflected in the API key's usage / credit consumption.
Actual Result
The two requests completed successfully but were not recorded in usage metering.
Additional Testing
Subsequent requests to the same model were metered normally:
| Request |
Usage |
| Streaming |
148 tokens / $0.0010 |
| Non-streaming |
109 tokens / $0.0007 |
| Long streaming |
12,415 tokens / $0.09 |
This indicates that usage metering works for both streaming and non-streaming requests to qwen/qwen3.7-max, but may intermittently fail for otherwise successful requests.
Impact
Successful inference requests may not be accounted for or billed, resulting in inaccurate usage reporting and credit consumption.
This is also relevant to the upcoming OpenRouter integration, where reliable per-request usage accounting is required.
Notes
Since subsequent requests to the same model were metered correctly, this appears to be an intermittent usage/metering ingestion issue, rather than a general streaming or model-specific billing failure.
Environment
qwen/qwen3.7-maxDescription
Two streaming requests to
qwen/qwen3.7-maxcompleted successfully, but no usage was recorded for either request.Both requests:
200finish_reason: "stop"[DONE]eventHowever:
Expected Result
Every successfully completed inference request should be recorded in usage metering and reflected in the API key's usage / credit consumption.
Actual Result
The two requests completed successfully but were not recorded in usage metering.
Additional Testing
Subsequent requests to the same model were metered normally:
This indicates that usage metering works for both streaming and non-streaming requests to
qwen/qwen3.7-max, but may intermittently fail for otherwise successful requests.Impact
Successful inference requests may not be accounted for or billed, resulting in inaccurate usage reporting and credit consumption.
This is also relevant to the upcoming OpenRouter integration, where reliable per-request usage accounting is required.
Notes
Since subsequent requests to the same model were metered correctly, this appears to be an intermittent usage/metering ingestion issue, rather than a general streaming or model-specific billing failure.