Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Guilherme Monteiro
guicybercode
·
AI & ML interests
smol-llms, llm-training, chinese-history, brazilian-history, production-ml, terminal-ui, data-curation, edge-ai
Recent Activity
updated a collection about 1 month ago
Culture, Mathematics, Technology & Ethics updated a Space about 1 month ago
guicybercode/culture-compass published a Space about 1 month ago
guicybercode/culture-compassOrganizations
None yet
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 1.81M • 242 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.85M • 430 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 411k • 136 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 275k • 225
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 66
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.48M • • 461 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 8.71M • • 653 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.42M • • 863 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 286k • 167
Culture, Mathematics, Technology & Ethics
Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 66
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.48M • • 461 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 8.71M • • 653 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.42M • • 863 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 286k • 167
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 1.81M • 242 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.85M • 430 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 411k • 136 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 275k • 225