Instructions to use Lapisbird/Llama-adaLR-model-latent-6-by-1_rl with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Lapisbird/Llama-adaLR-model-latent-6-by-1_rl with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("Lapisbird/Llama-adaLR-model-latent-6-by-1_rl") model = AutoModel.from_pretrained("Lapisbird/Llama-adaLR-model-latent-6-by-1_rl", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Improve model card: Add pipeline tag, paper link, GitHub link, and sample usage
Browse filesThis PR enhances the model card by:
- Adding the `pipeline_tag: text-generation` to improve discoverability on the Hugging Face Hub.
- Including a direct link to the paper [Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning](https://huggingface.co/papers/2511.21581).
- Providing a link to the official GitHub repository: https://github.com/apning/adaptive-latent-reasoning.
- Adding a sample usage code snippet from the GitHub README to guide users on how to load and use the model.
Please review and merge this PR if everything looks good.
README.md
CHANGED
|
@@ -1,10 +1,32 @@
|
|
| 1 |
---
|
| 2 |
-
library_name: transformers
|
| 3 |
-
license: llama3.2
|
| 4 |
base_model: meta-llama/Llama-3.2-1B-Instruct
|
| 5 |
datasets:
|
| 6 |
- whynlp/gsm8k-aug
|
|
|
|
|
|
|
| 7 |
tags: []
|
|
|
|
| 8 |
---
|
| 9 |
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
base_model: meta-llama/Llama-3.2-1B-Instruct
|
| 3 |
datasets:
|
| 4 |
- whynlp/gsm8k-aug
|
| 5 |
+
library_name: transformers
|
| 6 |
+
license: llama3.2
|
| 7 |
tags: []
|
| 8 |
+
pipeline_tag: text-generation
|
| 9 |
---
|
| 10 |
|
| 11 |
+
# Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning
|
| 12 |
+
|
| 13 |
+
This model introduces adaptive-length latent reasoning, a novel approach that uses a post-SFT reinforcement-learning methodology to optimize reasoning length while maintaining accuracy. It is presented in the paper: [Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning](https://huggingface.co/papers/2511.21581).
|
| 14 |
+
|
| 15 |
+
The official PyTorch implementation and training scripts are available on GitHub: https://github.com/apning/adaptive-latent-reasoning.
|
| 16 |
+
|
| 17 |
+
## Sample Usage
|
| 18 |
+
|
| 19 |
+
You can load these models using the function `automodelforcausallm_from_pretrained_latent` from `src.model_creation` as shown below.
|
| 20 |
+
|
| 21 |
+
```python
|
| 22 |
+
from transformers import AutoTokenizer
|
| 23 |
+
from src.model_creation import automodelforcausallm_from_pretrained_latent
|
| 24 |
+
|
| 25 |
+
repo_id = "Lapisbird/Llama-adaLR-model-latent-6" # Example model from the paper
|
| 26 |
+
|
| 27 |
+
model = automodelforcausallm_from_pretrained_latent(repo_id)
|
| 28 |
+
tokenizer = AutoTokenizer.from_pretrained(repo_id)
|
| 29 |
+
|
| 30 |
+
# Example usage with the loaded model (you would typically use this for text generation)
|
| 31 |
+
# For full inference examples, refer to the GitHub repository.
|
| 32 |
+
```
|