→ ~50× smaller than the 1M-parameter TinyStories model → ~3,000× smaller than AlexNet → 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️