-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 222 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
Collections
Discover the best community collections!
Collections including paper arxiv:2507.20534
-
moonshotai/Kimi-K2-Thinking
Text Generation • 1T • Updated • 42.3k • • 1.71k -
moonshotai/Kimi-K2-Instruct-0905
Text Generation • 1T • Updated • 37.7k • • 786 -
moonshotai/Kimi-K2-Instruct
Text Generation • 1T • Updated • 164k • • 2.38k -
moonshotai/Kimi-K2-Base
Text Generation • 1T • Updated • 14.1k • 306
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 222 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
-
moonshotai/Kimi-K2-Thinking
Text Generation • 1T • Updated • 42.3k • • 1.71k -
moonshotai/Kimi-K2-Instruct-0905
Text Generation • 1T • Updated • 37.7k • • 786 -
moonshotai/Kimi-K2-Instruct
Text Generation • 1T • Updated • 164k • • 2.38k -
moonshotai/Kimi-K2-Base
Text Generation • 1T • Updated • 14.1k • 306