Instructions to use froggeric/Qwen-Fixed-Chat-Templates with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use froggeric/Qwen-Fixed-Chat-Templates with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen-Fixed-Chat-Templates froggeric/Qwen-Fixed-Chat-Templates
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Keeps emitting <tool_call> json into the context
Ever since updating to merged 3.5/6 version randomly stops and emits json for tool_call
Using unsloth/Qwen3.6-27B-MTP-GGUF:Q5_K_XL in llama-server with q8 k/v cache.
<tool_call>
<function=bash>
<parameter=command>
find Library/PackageCache -name "UniversalResourceData.cs" 2>/dev/null
</parameter>
</function>
</tool_call>
Although this is in pi so I wonder if something broke there 🙃
Looks like it believed it was a thinking block
{
"type": "message",
"id": "ccbda428",
"parentId": "3318fa1e",
"timestamp": "2026-05-18T16:32:49.150Z",
"message": {
"role": "assistant",
"content": [
{
"type": "thinking",
"thinking": "<tool_call>\n<function=bash>\n<parameter=command>\nfind Library/PackageCache -name \"UniversalResourceData.cs\" 2>/dev/null\n</parameter>\n</function>\n</tool_call>",
"thinkingSignature": "reasoning_content"
}
],
"api": "openai-completions",
"provider": "llama-cpp-mac",
"model": "Qwen3.6-27B",
"usage": {
"input": 671,
"output": 43,
"cacheRead": 58402,
"cacheWrite": 0,
"totalTokens": 59116,
"cost": {
"input": 0,
"output": 0,
"cacheRead": 0,
"cacheWrite": 0,
"total": 0
}
},
"stopReason": "stop",
"timestamp": 1779121949860,
"responseId": "chatcmpl-xVLlDQ1nqLjtHoIgiuaBwOcwrNYRto06"
}
}
Hi @theliphant , if you want you give my just released template rebuild project a try.
https://huggingface.co/StableQuant/Qwen-Templates-Rebuild-Project/
If v1.0 doesnt fixes it for you please tell me. If it works tell me too :-)
It will then be added into the research pipeline
Regards
In v20, I rewrote the tool formatting instructions to be much stricter about keeping XML structures and avoiding JSON drift. Upgrading to the latest template should fix this behavior.
I just checked, v20 still happens then stops.
<tool_call>
<function=bash>
.....
</function>
</tool_call>
💀
spiritbuun template isn't affected, however It exhibits phenomenon of repeatedly calling duplicate tools when using Chrome DevTools on the Pi on both your new template and bun.
Taming this thing is really annoying. Qwen want us to pay for best.
Yes, unfortunately as this is not enforced algorithmically but through a system prompt, it is not 100% reliable.
I'm not sure:
auto_disable_thinking_with_tools = false will cause a tool_call 💀
auto_disable_thinking_with_tools = true will not cause an error but sometimes causes a duplicate.
Hi! The massive v21 update was just released. It completely overhauled tool-calling compatibility (switching to native Hermes JSON for inference engines like llama.cpp and LM Studio), fixed the preserve_thinking amnesia stalls/loops, and resolved several </think> parsing and prompt injection bugs.
This issue should now be fully resolved in v21. I am closing this discussion, but please feel free to reopen or create a new one if you are still experiencing any trouble. Thanks!
Update: We actually reverted the JSON switch in v21.1 because it critically broke native vLLM parsers out of the box (like qwen3_coder). Qwen models are specifically trained on the XML <function=...> syntax, so sticking to the native XML format is necessary for universal engine support.