Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
25.0
TFLOPS
Andrew DeLisa
ayan4m1
14
13
Follow
aiqualitylab's profile picture
John6666's profile picture
colinurbs's profile picture
10 followers
·
9 following
https://andrewdelisa.com
ayan4m1
AI & ML interests
Distilled fine-tuning
Recent Activity
new
activity
about 12 hours ago
mradermacher/model_requests:
ayan4m1/Qwen2.5-Coder-14B-E2E-Tests
updated
a model
about 14 hours ago
ayan4m1/Qwen2.5-Coder-14B-E2E-Tests
reacted
to
Felladrin
's
post
with 🔥
about 15 hours ago
I've open-sourced the trainer I've been using to build tiny language models from scratch, together with the 95M base model I trained with it. The trainer runs on Deno (https://deno.com, cross-platform), trains on WebGPU, and it writes GGUF directly. No Python/PyTorch. The weights live in a GGUF file from the first step to the last, so every checkpoint is already something llama.cpp can load. The model is https://huggingface.co/Felladrin/Minueza-3-95M-Base: 94.7M parameters, 1.95B tokens seen, 8192 context. And here’s the repository on GitHub: https://github.com/felladrin/gguf-trainer Here on Hugging Face, I published the optimizer state next to the weights, so you can continue the pretraining instead of starting over. Or start your own from nothing: `deno run -A cli.ts demo` trains a tiny one end to end in under a minute. And the docs are written for coding agents, so you can point your agent of choice at the GitHub repo and have it drive the whole pipeline.
View all activity
Organizations
ayan4m1
's Spaces
1
Sort: Recently updated
Runtime error
Agents
Prompt Enhancer
💬
Enhance image generation prompts