Replies: 6 comments 4 replies
|
this looks close to a couple existing llama.cpp threads rather than a totally new config issue. #21445 has the accepted pointer for per-request control: For the Step 3.7-specific loop/overthinking behavior with tool-ish reasoning, #24181 is probably the better thread to watch/add your repro details to. |
|
I don't think Step 3.7 Flash really supports reasoning efforts like GPT-OSS. Even though it is documented, I've tried setting it to "low" and saw no difference in its output when using the official API. In the end the amount of reasoning seems to be decided by the complexity of the task and some randomness. llama.cpp reasoning budget options can work quite well if set up proper.y Before I reported the parser bug (#24181), I had been using reasoning_budget as a workaround. See this for more details: https://huggingface.co/stepfun-ai/Step-3.7-Flash-GGUF/discussions/6 |
|
Interesting discussion! Optimizing reasoning behavior is important for creating AI systems that are both efficient and responsive. This is especially valuable for food-related applications, where AI can recommend recipes, suggest restaurant menus, or provide nutrition guidance without unnecessary delays. Better control over reasoning effort leads to a smoother experience for users. |
|
Thanks for bringing attention to this issue. Fine-tuning how AI handles reasoning effort can improve performance while reducing unnecessary computation. In the food industry, this could help AI-powered meal planners, recipe assistants, and menu recommendation systems deliver faster, more accurate results, making healthy and delicious choices easier for everyone. |
|
This is an interesting discussion about how Step AI 3.7 handles reasoning behavior and configuration settings. Reports like this are valuable because they help developers identify inconsistencies and improve the model's performance in future updates. On a different note, I recently found a helpful guide on Necklace Length for Kids that explains how to choose safe and comfortable necklace lengths for children of different ages and styles. |
|
First, I just wanted to say this is an awesome project: it significantly improved the speed when using Qwen3.8-27B on my RTX3090! This fork of llama.cpp does not seem to recognize: "reasoning-effort": So, I'm using (command-line): Or (config.ini): |
Uh oh!
There was an error while loading. Please reload this page.
I wonder if I'm doing something wrong. The 3.7 model seems to massively overthink even simple questions such as writing C++ AVL implementation. It's to the point where it's effectively several times slower than comparably-sized models. It does not get into a loop and eventually finishes.
However, I don't see the value of "checking" these test cases when it's actually NOT executing the code and testing. It's just regurgitating test cases.
I was planning to use this for non-coding scenarios but coding is simpler to test. Thanks.
Arguments: --jinja --chat-template-kwargs {"reasoning_effort":"low"}
Llama.cpp: b9496
Model: bartowski/Step-3.7-Flash-GGUF
Quant: Q8_0
MTP layer: Step3.7-flash-mtp-Q8_0.gguf
I observed the same behavior with the stepfun-ai/Step-3.7-Flash-GGUF model. MTP is working and t/s increases by 25% from 8 t/s to 10 t/s.
Code Thinking Snippet:
Test Case Thinking Snippet:
All reactions