WorkflowEvals Collection https://github.com/typesafe-ai/WorkflowEvals • https://evals.typesafe.ai • 4 items • Updated about 18 hours ago • 8
MiMo-V2.6-RL in Harbor Collection All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness. • 8 items • Updated 3 days ago • 6
K2 Horizon Collection K2 Horizon models, datasets, and supporting resources • 24 items • Updated about 18 hours ago • 138
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU Paper • 2604.05091 • Published Apr 6 • 45
NVIDIA Nemotron v3 Collection Open, Production-ready Enterprise Models • 33 items • Updated Aug 14 • 378
pplx-embed Collection Diffusion-Pretrained Dense and Contextual Embeddings • 12 items • Updated 26 days ago • 103
Embeddings datasets ⚡️ Collection This collection gather datasets for embeddings pre-training and fine-tuning. • 19 items • Updated 9 days ago • 5
SWE-bench Collection SWE-bench is a benchmark for evaluating Language Models and AI Systems on their ability resolve real world GitHub Issues. • 4 items • Updated Mar 8, 2025 • 10
MMFineReason Collection High-quality STEM reasoning dataset for Multimodal LLM post-training. • 8 items • Updated May 7 • 24
view article Article The Optimal Architecture for Small Language Models codelion • Dec 26, 2025 • 123
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Paper • 2512.07461 • Published Dec 8, 2025 • 80
Ministral 3 Collection A collection of edge models, with Base, Instruct and Reasoning variants, in 3 different sizes: 3B, 8B and 14B. All with vision capabilities. • 9 items • Updated Dec 2, 2025 • 173