Image-Text-to-Text
Safetensors
English
qwen3_5_moe
qwen
qwen3.6
Mixture of Experts
bf16
speculative-decoding
speculators
dflash
multimodal
vision
conversational
Instructions to use ericpandev/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
Important
If you downloaded the model before May 9, please re-download the model as some tensors were not in the right order
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-bf16
Bf16 safetensors conversion of original GGUF model + mmproj, in exact HuggingFace format matching Qwen/Qwen3.6-35B-A3B. Ready for vLLM speculators (DFlash/EAGLE-3) as a verifier/target model.
Original Model
- Creator: HauhauCS
- Source: Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
- Architecture: Qwen3_5MoeForConditionalGeneration (multimodal)
- Base: Qwen 3.6 35B MoE (256 experts, 8 active, 3B active params)
Conversion
- Text model: 733 GGUF Q8_K_P tensors → dequantized to bf16
- Vision encoder: 334 GGUF F16 tensors → merged (patch_embd slices combined)
- Full tokenizer extracted (248,320 vocab + 247,587 BPE merges + chat template)
- 1026 HF weights (693 text + 333 vision) across 17 shards, 70.2 GB total
- All weight names match reference model exactly (e.g.
model.language_model.*,model.visual.*)
Specs
- Layers: 40 (30 linear attention/SSM + 10 full attention every 4th)
- Hidden: 2048, Heads: 16 Q / 2 KV, Head dim: 256
- Linear attention: key/value head dim 128, conv kernel 4
- Experts: 256 routed + 1 shared per layer, 8 active
- Vision: 27-layer encoder, Qwen3VL merger, patch 16, 768 image size
- Vocab: 248,320, Context: 262,144
- Downloads last month
- 806