nielsr HF Staff commited on
Commit
a67dcd9
·
verified ·
1 Parent(s): a678e0e

Improve model card with paper, code, pipeline tag, and usage

Browse files

This PR enhances the model card by:
- Linking it to the paper [Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning](https://huggingface.co/papers/2511.21581).
- Adding a link to the GitHub repository (https://github.com/apning/adaptive-latent-reasoning).
- Incorporating the `pipeline_tag: text-generation` to ensure discoverability on the Hub.
- Expanding the model description and including a sample usage snippet for loading the model, extracted directly from the GitHub README.

Please review and merge if these improvements are satisfactory.

Files changed (1) hide show
  1. README.md +24 -3
README.md CHANGED
@@ -1,10 +1,31 @@
1
  ---
2
- library_name: transformers
3
- license: llama3.2
4
  base_model: meta-llama/Llama-3.2-1B-Instruct
5
  datasets:
6
  - whynlp/gsm8k-aug
 
 
 
7
  tags: []
8
  ---
9
 
10
- Built with Llama
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
 
 
2
  base_model: meta-llama/Llama-3.2-1B-Instruct
3
  datasets:
4
  - whynlp/gsm8k-aug
5
+ library_name: transformers
6
+ license: llama3.2
7
+ pipeline_tag: text-generation
8
  tags: []
9
  ---
10
 
11
+ # Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning
12
+
13
+ This repository contains the model weights associated with the paper [Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning](https://huggingface.co/papers/2511.21581).
14
+
15
+ Latent reasoning represents a new development in Transformer language models that has shown potential in compressing reasoning lengths compared to chain-of-thought reasoning. This work introduces adaptive-length latent reasoning models and a post-SFT reinforcement-learning methodology to optimize latent reasoning length by minimizing it while maintaining accuracy. This approach further reduces compute usage and raises the bar on the compressive capabilities of latent reasoning models. Experiments on the Llama 3.2 1B model and the GSM8K-Aug dataset demonstrated a 52% drop in total reasoning length with no penalty to accuracy.
16
+
17
+ For more details and the full code, please refer to the [GitHub repository](https://github.com/apning/adaptive-latent-reasoning).
18
+
19
+ ## Usage
20
+
21
+ You can load these models using the `transformers` library with a custom function provided in the project's `src.model_creation`. An example is provided below:
22
+
23
+ ```python
24
+ from transformers import AutoTokenizer
25
+ from src.model_creation import automodelforcausallm_from_pretrained_latent
26
+
27
+ repo_id = "Lapisbird/Llama-adaLR-model-latent-6" # Replace with the specific model you want to load from the paper's collection
28
+
29
+ model = automodelforcausallm_from_pretrained_latent(repo_id)
30
+ tokenizer = AutoTokenizer.from_pretrained(repo_id)
31
+ ```