123
followers ยท
450 following AI & ML interests Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity reacted to comgen42 's post with ๐ฅ about 8 hours ago Kodiak v0.4 is out: cortex-agent-llc/kodiak-v0.4-1b (plus an accuracy mode that averages three runs).
This release came out of public feedback. After v0.3, someone showed that two of its skills were answering from keywords instead of reading the case. So we built tests a keyword shortcut can't pass: the same case twice, with one detail changed so the right answer flips.
Refund eligibility is now fixed: it gets both versions right 79% of the time, up from 22%. The push-to-main-with-failing-tests case that started this now gets "ask the user first". Agent step safety improved but isn't fixed, and the model card says plainly not to use it as a safety control. Sarcasm still shows no signal on real tweets, and accuracy on brand-new kinds of task hasn't moved. That's the big problem for the next version.
Now a pause. I'm shutting the training box down for two weeks while I travel in Turkey. I'll be walking through ancient ruins: places where people built things that lasted thousands of years without any of our tools. I'm hoping that does what travel usually does for me, shakes loose some ideas. I'll be thinking about how to teach Kodiak to handle tasks it has never seen, how to run my publishing company better, and where agentic automation actually earns its keep.
No training while I'm gone. When I'm back, Kodiak gets the ideas.
Every number, including the misses, is in the public build log: github.com/grizzlypeaksoftware/kodiak View all activity Organizations