Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
omar81939
's Collections
DIAL: GRPO on looped models
RL4RLM: Training Native Recursive Language Models
DIAL: GRPO on looped models
updated
Aug 13
Random-depth SFT and depth-as-action GRPO checkpoints for Ouro-1.4B-Thinking.
Upvote
1
Sort: Collection
omar81939/Ouro-1.4B-Thinking-depth-SFT
Text Generation
•
1B
•
Updated
Aug 13
•
29
omar81939/Ouro-1.4B-Thinking-depth-GRPO
Text Generation
•
1B
•
Updated
Aug 13
•
36
Upvote
1
Sort: Collection
Share collection
View history
Collection guide
Browse collections