- GLM 5.2 - Flux 3 - New Qwen model - New small model leaderboards - Lots of people finetuning smol models. - Some even under 12 year olds clauders are here (was not on my bingo card this year) - ChatGPT's Sol became a lot faster this week - LFM2.5 2.6b - Kimi K3 (though only a few will run it) - New Ling 3.0 Tiny - New video model that is making south park videos? - Deepseek v4 flash being more honest than bigger models - The new model from meta
bench-labs developed **GCTokenizer-v1**, which is a multi-lingual tokenizer Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL" It utilizes an encoding scheme which allows it to handle characters in any language around the world
General (multi lingual) Consensus (from multiple model tokenizers consensus) Tokenizer
We included an implementation script too, built like BPE- it can encode arbitrary text, most of the time, efficiently
Our analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. ๐ง
Today, we are announcing a brand-new series of SupraLabs models: Supra2 This series will feature various models, including such as: - ๐ Supra2-Nano (0.4M) โ The smallest Supra2 model. - ๐ค Supra2-Small (1.4M) โ The tiny model that runs everywhere. - ๐ช Supra2-Medium (25M) โ Our medium class model in the Supra2 family. The powerful midsizer. - ๐ฅ Supra2-Pro (100M): base, instruct, reasoning, code, math and more! โ The most capable model yet! A real allrounder for all your everyday tasks. - ๐จ Supra2-IMG โ our generative text-to-image model ...and many more...
Current progress: - Nano (0.4M) and Small (1.4M): in training; almost done. Baseline set. - Medium (25M): coming soon... - Pro (100M): in training; finishes in 66 hours - Monday, 3rd August 2026, 12:00AM - IMG: coming soon...
You can support us with a like and follow if you want! Don't miss our next release! Stay tuned...
small reasoning models are overrated, these little ones just doom loop a lot by default. good data will always be the moat when training or finetuning small models and latest sota models like fable 5 and gpt 5.6 are increasingly making this a lot easier to do.