Qwen3-32B-FP8
This repository contains a pre-compiled build of Qwen/Qwen3-32B-FP8 for running it on FuriosaAI RNGD with Furiosa-LLM.
Overview
Qwen3-32B is the dense 32B model in the Qwen3 series, an auto-regressive transformer that supports both a thinking mode for complex reasoning and a non-thinking mode for general dialogue, along with strong instruction following, multilingual coverage, and tool usage. Its intended use is the same as the upstream Qwen/Qwen3-32B-FP8, and it is released under the Apache 2.0 License.
- Architecture: Qwen3 (dense)
- Input / Output: Text / Text
- Supported Inference Engine: Furiosa LLM
- Supported Hardware: FuriosaAI RNGD
Quantization
Weights are quantized to FP8 (static), following the upstream FP8 release, and activations use dynamic FP8 quantization at runtime (per-token / per-block). The KV cache stays in 16-bit precision.
Features
- Reasoning. Qwen3-32B is a hybrid reasoning model: thinking mode is on by default and can be toggled per request via
enable_thinking. Launch the server with--reasoning-parser qwen3to have the reasoning content returned in a separate field (see Basic Usage below). - Tool calling. The model supports tool (function) calling through the
hermestool-call parser, the parser used by the Qwen3 series.
Parallelism Strategy
On RNGD, Qwen3-32B-FP8 runs with a tensor-parallel size of 32 PEs, which maps to four RNGD cards (8 PEs per card).
Usage
To run this model with Furiosa-LLM, follow the example commands below after installing Furiosa-LLM and its prerequisites.
Launch the server
Serve the model with the qwen3 reasoning parser so the chain of thought is
returned in a separate field:
furiosa-llm serve furiosa-ai/Qwen3-32B-FP8 \
--reasoning-parser qwen3
To also enable tool (function) calling, add the hermes tool-call parser (the
parser used by the Qwen3 series):
furiosa-llm serve furiosa-ai/Qwen3-32B-FP8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser hermes
When the server is ready, you will see:
INFO: Started server process [27507]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
Basic Usage
The server exposes an OpenAI-compatible API. You can send a request with curl:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "furiosa-ai/Qwen3-32B-FP8",
"messages": [{"role": "user", "content": "What is the capital of France?"}]
}' \
| python -m json.tool
With --reasoning-parser qwen3, the thinking content is returned separately from
the final answer:
response.choices[].message.reasoning(non-streaming)response.choices[].delta.reasoning(streaming)
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="furiosa-ai/Qwen3-32B-FP8",
messages=[{"role": "user", "content": "How many r's are in 'strawberry'?"}],
)
print("Reasoning:", response.choices[0].message.reasoning)
print("Answer:", response.choices[0].message.content)
Note: The
reasoningfield is not part of the OpenAI API specification but is a widely followed convention (the OpenAI Agents SDK, vLLM, and others). It appears only in responses that contain reasoning content; accessing it otherwise raises anAttributeError.
Advanced Usage
Toggling thinking. Qwen3-32B is a hybrid model: it reasons by default and can
switch thinking on and off. To turn thinking off for a single request, pass
enable_thinking through chat_template_kwargs; the response then carries no
reasoning content, so read only message.content:
# Disable thinking for a single request
response = client.chat.completions.create(
model="furiosa-ai/Qwen3-32B-FP8",
messages=[{"role": "user", "content": "What is the capital of France?"}],
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)
To default every request to non-thinking, launch the server with
--default-chat-template-kwargs (a request can still re-enable thinking with its
own chat_template_kwargs):
furiosa-llm serve furiosa-ai/Qwen3-32B-FP8 \
--reasoning-parser qwen3 \
--default-chat-template-kwargs '{"enable_thinking": false}'
Tool calling. With the server launched using
--enable-auto-tool-choice --tool-call-parser hermes (see
Launch the server), pass tools in the request and let the
model decide when to call them. See the
Tool Calling guide
for a complete client example and details on tool-choice options.
Learn more
- Tool Calling โ parsers, tool-choice options, and more examples
- Furiosa-LLM Server (
furiosa-llm serve) โ full OpenAI-compatible API reference and serving options - Qwen/Qwen3-32B-FP8 โ upstream model card
- Downloads last month
- 1,706