DavidAU commited on
Commit
c03dd4f
·
verified ·
1 Parent(s): ea3e95a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +41 -0
README.md CHANGED
@@ -44,6 +44,10 @@ base_model:
44
  - DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
45
  ---
46
 
 
 
 
 
47
  <I>The Qwen 3.5 version (also 40B) got 181 likes+ This version uses the new Qwen 3.6 27B arch (which exceeds even Qwen's own 398B model).</I>
48
 
49
  <small><b><font color="red">WARNING:</font></B> This model has character and intelligence. It will take no prisoners. It will give no quarter. Uncensored,
@@ -90,6 +94,43 @@ in the third act because you can't see the plot holes you've been digging since
90
 
91
  ---
92
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
93
  <B>Qwen 3.5 40B Version: 181 likes and counting... </B>
94
 
95
  https://huggingface.co/DavidAU/Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
 
44
  - DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
45
  ---
46
 
47
+ [ building / uploading... ]
48
+
49
+ <small><font color="red">Ultimate NEO GGUF QUANTS:</font>Custom built DUAL Imatrix NEO-CODER quants that exceed all other quants in terms of quality, stability, precision and long convo usage.</small>
50
+
51
  <I>The Qwen 3.5 version (also 40B) got 181 likes+ This version uses the new Qwen 3.6 27B arch (which exceeds even Qwen's own 398B model).</I>
52
 
53
  <small><b><font color="red">WARNING:</font></B> This model has character and intelligence. It will take no prisoners. It will give no quarter. Uncensored,
 
94
 
95
  ---
96
 
97
+ <B>NEO-CODE-Di-IMatrix-MAX-GGUF Quants:</B>
98
+
99
+ Quant "engineering" focused on balance and precision, vs raw power (which seemed in some cases to destabilize the model/quant).
100
+
101
+ In other words benchmarks / stats determined the best quants, not guesswork or one size fits all approach.
102
+
103
+ This was done to ensure long context, long/multi-convos, coding and math etc etc performed as close as possible to full precision model as well as one-shot, and standard prompting / problem solving.
104
+
105
+ TWO Imatrix datasets were used to do this by first getting "raw stats" on both, then merging them to get the best of each imatrix in one dataset then this was used to make the "NEO-CODE-Di-IMatrix-MAX" quants.
106
+
107
+ Additional tensor adjustments were also made, which were also measured (benched) and adjusted too.
108
+
109
+ <B>GGUF POWER UPS:</B>
110
+
111
+ A radically stronger, more potent GGUF for all use cases.
112
+
113
+ Meets Unsloth quality, and exceeds it in some metrics (see below).
114
+
115
+ DETAILS:
116
+ - DI-MATRIX (duel imatrix) of NEO and NEO-CODE imatrix datasets (by DavidAU).
117
+ - All Unsloth tensor enhancements + additional enhancements CALIBRATED thru metrics testing.
118
+ - Every quant benchmarked against BF16/full precision model.
119
+ - There is a special Q8_0 quant, with BF16 components. Imatrix has no effect on Q8/BF16 tensors.
120
+
121
+ <B>VISION:</B>
122
+ - Vision (images) tested.
123
+ - You need an "mmproj" (just one) of these downloaded too, and placed in the same folder as the GGUF for images.
124
+
125
+ <B>Qwen Model Settings (suggested):</B>
126
+
127
+ - Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
128
+ - Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
129
+ - Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
130
+ - Context window min from 8k to 16k.
131
+
132
+ ---
133
+
134
  <B>Qwen 3.5 40B Version: 181 likes and counting... </B>
135
 
136
  https://huggingface.co/DavidAU/Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking