The fire of multiple Fable Fusion 711 cores in one, larger model.
This is the first 40B fine tune that reaches "closed source" (IE OpenAI, Claude) level of intelligence in both 8 bit and 4 bit.
This model is composed from multiple Qwen 27B Fable Fusion 711 cores (1700+ likes, 2.3 million+ downloads) - a record breaking model in terms of intelligence and raw power.
The "40B Eleanor" takes this to the next level with improvements in thinking tokens/ thinking block size (1/10 to 1/2 the size), thinking in general and output detail quality with deep analytics too.
- 1/10 to 1/2 the number of thinking tokens. - Extreme depth of detail in generations, including long form, in depth analytics. - STRONG creative abilities. - Auto-variable reasoning: Model only reasons as much as the task requires. - Strong general intelligence. - It says what it means in less words, more clearly than any previous tuned model. - It will go all in, in exacting detail when the situation calls for it. - If it thinks something is wrong / wrong path it will say so too.
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget.
Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through.
The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting.
DeepSeek plans to raise token prices. I don't think this is because they are bleeding, but they are overwhelmed. If your price is 1/10 of your affiliate vendors, you can't leverage their resources. Markup is the only way to diverge traffic away.
Sadly I haven't found discussions on differentiators enabling DS to balance cost at such low prices. All software solutions (that we know of) are accessible by other vendors. If you attribute it to electricity or hardware, you can't explain why GLM and Kimi charge so much for their APIs.
This is where our attention should be (but distracted by things above).