Is this model reliable for serious work?

#7
by Nunodonato - opened

Background: we are using qwen3.6-27B for important internal workflows at our company, as well as coding model for claude code.

I've been "shopping" around for possible upgrades without a big leap in minimum VRAM and bumped into this model. Unfortunately, without published benchmarks its hard to assess how good it might be and if it has any disadvantages compared to the original 27B.

Needless to say that the uncensored part is irrelevant for us as we don't deal with any kind of content that needs uncensoring. Its more about agentic and reasoning capabilities.

Any thoughts?

From my use, baseline 3.6 27B vs deckard 40B, almost indistinguishable on coding work.
GPQA - Diamond they scored within a half point on my tests, no difference worth mentioning.
Agentic tool calling and task completions, basically identical.

The one outlier I experienced is that 40B will occasionally reach for a more complex solution a given problem than the 27B did, and it doesn't always get that more elaborate solution right. Like it has the reasoning to see there's a better solution, but not the capacity to realize it fully. This only happened once or twice on pretty out there stuff almost certainly not in distribution tasks.

Where Deckard differs...the instruction following and adoption of a "persona" when given one, might have unexpected side effects based on various user prompts in work like...
"You are a senior software engineers do ...." or "You are an HR expert evaluate this for ...." whatever, the character flavor added might improve and might degrade performance, worth test, I quite like this model for general use.

Owner

The Deckard dataset is specifically for creative, however it will bump up some metrics [and have a slightly neg effect on others].
However the expansion from 27B to 40B may affect other parts of the model all by itself.

For programming / agentic -> use min Q6 (for all qwens) ; otherwise you will get some breakage/issues.

NOTE: a lot of issues with agentic can be traced to issue(s) in the jinja template(s).

You could prove by an example, where deckard 40b q6 > qwen 27b q8 mtp.

I find that the 40B (Q8) seem to loose track and eventually collapse into nonsense output in my creative writing tests far more often then 27B. But it might also be my settings. I am not exactly a pro at this, but I thought I should mention it in case other people have noticed something similar.

*Edit

My apologies, I was referring to the Qwen3.6-40B-Deck-Opus-NEO-CODE-HERE-2T-OT-HIGH-Q8_0.gguf

Sign up or log in to comment