Instructions to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16
- SGLang
How to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16", max_seq_length=2048, ) - Docker Model Runner
How to use Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16
PINQWEN-3.6-35B-CLEAN-BF16
Full 16-bit merged weights of PINQWEN-3.6-35B-CLEAN, a supervised fine-tune of Qwen/Qwen3.6-35B-A3B by Blackfrost-AI. General-purpose assistant tuned for reasoning and agentic tool-use. Safety-aligned (CLEAN) variant.
Model Description
PINQWEN-3.6-35B-CLEAN is a supervised fine-tune (SFT) of Alibaba's Qwen/Qwen3.6-35B-A3B base model. This repository holds the BF16 release: the full 16-bit merged weights.
- Developer: Blackfrost-AI
- Base model: Qwen/Qwen3.6-35B-A3B (Alibaba Qwen team)
- Architecture:
qwen3_5_moe(Qwen3_5MoeForConditionalGeneration) — a Mixture-of-Experts model with 256 experts (8 routed + 1 shared per token), ~36.97B total parameters and ~3B active per token. Hybrid Gated-DeltaNet + gated attention. Unified vision-language model with 262K native context. Thinking-on by default. - Variant: CLEAN — the aligned variant. This is not an abliterated model; the base model's safety alignment is preserved.
- License: Apache-2.0
- Language/modality: Text (see Limitations regarding vision).
A 4-bit NVFP4 quantization of this same model is released separately at Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-NVFP4.
PINQWEN & The Void
PINQWEN is Blackfrost-AI's model name for this Qwen3.6-based series. The CLEAN suffix marks the safety-aligned build.
The model was fine-tuned on The Void (v4), Blackfrost's proprietary distillation corpus: roughly 5,032 high-quality multi-turn examples that blend distilled reasoning / chain-of-thought and broad knowledge with agentic, ReAct-style tool-use trajectories drawn from multiple frontier teacher models. Refusal and denial data is scrubbed from the corpus to protect MoE routing quality. The corpus construction method is proprietary and not disclosed here.
Training Procedure
- Method: Supervised fine-tuning (SFT) with bf16 LoRA via Unsloth. LoRA was run in bf16 rather than 4-bit — 4-bit QLoRA degrades this MoE.
- LoRA configuration: rank 32, alpha 32, dropout 0. Adapters were applied to the attention projections (q/k/v/o) and the MoE expert projections (
gate_up_proj,down_proj), giving 1.86B trainable parameters (5.04% of the model). - Schedule: 3 epochs, learning rate 2e-4, linear schedule, length-grouped batching.
- Objective: Multi-turn SFT with loss masked to assistant turns only.
- Hardware: 8× NVIDIA B200, DDP.
- Release: The LoRA adapter was merged back to 16-bit for this release.
Vision (multimodal)
This is a vision-language model. It carries a full vision tower inherited from the Qwen3.6-35B vision-language base, so it accepts images and video alongside text.
Scope of Blackfrost's work: training here was text-only; the vision tower is inherited unchanged from the base and was not tuned or evaluated by Blackfrost. Multimodal behavior tracks the base model — validate it for your use case.
Intended Uses
General-purpose text assistant for reasoning and agentic / tool-use (ReAct-style) tasks.
Limitations
- Text-focused SFT. The base model is vision-language, but vision was not specifically tuned or evaluated in this work. Treat vision behavior as untuned base-model behavior.
- No public benchmark numbers are claimed yet. Internal evaluations are pending; no scores are reported here.
- The model may inherit base-model limitations and can hallucinate. Verify important facts.
How to Use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [
{"role": "user", "content": "Explain what a Mixture-of-Experts model is in two sentences."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
License & Attribution
Released under Apache-2.0. Built on Qwen/Qwen3.6-35B-A3B by Alibaba's Qwen team (Apache-2.0). Fine-tuning and release by Blackfrost-AI.
Responsible Use
This model retains the base model's safety alignment. Do not use it to generate content that exploits minors or promotes self-harm. Standard responsible-use expectations apply.
- Downloads last month
- 986
