Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Guilherme Monteiro
guicybercode
·
AI & ML interests
smol-llms, llm-training, chinese-history, brazilian-history, production-ml, terminal-ui, data-curation, edge-ai
Recent Activity
updated a collection 12 days ago
Culture, Mathematics, Technology & Ethics updated a Space 12 days ago
guicybercode/culture-compass published a Space 12 days ago
guicybercode/culture-compassOrganizations
None yet
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.55M • 231 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.36M • 414 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 484k • 128 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 298k • 220
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.69M • • 443 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 6.43M • • 620 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.33M • • 822 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 190k • 148
Culture, Mathematics, Technology & Ethics
Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.69M • • 443 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 6.43M • • 620 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.33M • • 822 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 190k • 148
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.55M • 231 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.36M • 414 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 484k • 128 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 298k • 220