Instructions to use jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jucamohedano/rlm-laguna-xs2-dense-3b-v0.1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("jucamohedano/rlm-laguna-xs2-dense-3b-v0.1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jucamohedano/rlm-laguna-xs2-dense-3b-v0.1
- SGLang
How to use jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 with Docker Model Runner:
docker model run hf.co/jucamohedano/rlm-laguna-xs2-dense-3b-v0.1
RLM-Laguna-XS.2-Dense-3B · v0.1 (protocol SFT warm start)
A ~3.0 B dense model warm-started for the Recursive Language Model (RLM) protocol: the model emits Python inside a ```repl fence, the harness executes it against an offloaded long context, and the model reads the result and decides its next action.
This checkpoint is a partial result. It reliably produces and executes valid REPL actions and attempts submissions, but it solves only a minority of the tasks we measured. See Honest status before using it.
Credits and lineage
This model would not exist without the work of others. Please credit them if you build on it.
Base model — the Poolside Laguna MoE teacher
poolside/Laguna-XS.2— the 33.4 B-total / 3.0 B-active MoE (256 experts, top-8) from Poolside AI. This is the teacher model for the whole lineage.
The hackathon team — MoE-to-dense distillation and CUDA SFT
Everything between the MoE teacher and this checkpoint was produced by a hackathon team and published openly:
- Model author on Hugging Face:
EvanOLeary(recon, kernel-mix, and the-cuda-sft/-cuda-sft-v2checkpoints). - Recipe and code:
Tyronita/laguna-dense-cuda-kernels— the full end-to-end recipe: densification, reconstruction pretraining, SFT, GRPO, DPO, and the evaluation harnesses. - Provenance credited by them:
cm2435/laguna-xs2-expert-coactivation-scheduling.
What that team built, which our checkpoint inherits:
- MoE → dense densification (K=8) with a DO-ACP expert-selection warm start rather than random init, which was the single biggest convergence lever.
- Reconstruction pretraining of the dense FFN against the teacher's MoE-block activations (teacher-forced, all layers in parallel).
- SFT on CUDA kernel data (Sakana
AI-CUDA-Engineer-Archive, levels 1+2, then a follow-up SFT on levels 1+2+3) — that follow-up SFT is exactlyEvanOLeary/laguna-xs2-dense-k8-cuda-sft-v2, our base model.
The RLM paradigm
- Recursive Language Models, Alex Zhang et al. — arXiv:2512.24601,
blogpost, github.com/alexzhang13/rlm.
Our repository layout, environment and training stack come from that project, and the model-name
convention (
rlm-<base-slug>-v<major>.<minor>, cf.mit-oasys/rlm-qwen3-30b-a3b-v0.1) follows theirs. - prime-rl / verifiers (Prime Intellect) for the trainer and the multi-turn environment.
Datasets
- OOLONG long-context benchmark tasks (
oolongbench/oolong-synth) as the SFT/eval task family. - Teacher trajectories for SFT were distilled from a larger model on those tasks.
What this checkpoint adds
Starting from a CUDA-kernel SFT model, we teach the RLM action protocol. Concretely, the model
learns to open a ```repl fence, write short executable Python, close the fence, and — when it
believes it has the answer — set answer["ready"] = True.
Training: LoRA (r=32, α=64) on attention and FFN projections, 60 clean single-target-turn rows, resumed from an earlier adapter checkpoint. Best checkpoint by validation loss is step 50 of a 70-step run; later steps overfit.
Measured behaviour (greedy and T=0.7, real environment execution)
base (cuda-sft-v2) |
this checkpoint (step 50) | |
|---|---|---|
| emits a closed ```repl fence | 0 / 3 tasks | 3 / 3 tasks |
| REPL blocks actually executed | 0 | 3–6 per task |
| attempts a submission | 0 / 3 | 2 / 3 |
| correct answer (reward 1) | 0 / 3 | 1 / 3 |
A representative first action, produced verbatim:
print(type(context))
print(len(context))
print(context[:500])
The failure mode moved from format failure to reasoning failure: the model now produces valid, executed actions, and its remaining errors are ordinary coding mistakes (for example a regex that does not match the date format) plus a tendency to loop when it cannot make progress.
Honest status
Please read this before drawing conclusions.
- Reward is 1 on only 1 of 3 measured tasks. This is a warm start, not a finished RLM.
- The evaluation tasks are teacher-seen. Our SFT traces were distilled over the same task pool (103 OOLONG spam tasks under our filters, of which the teacher consumed all 103). The held-out rows are held out from the SFT split, not from the teacher. These numbers are not a generalization measurement.
- Validation loss is not a generalization number either, for the same reason. It fell 1.52 → 1.31 by step 50 and then rose to 1.62 by step 70 while training loss fell to 0.14, i.e. overfitting.
- Known weaknesses: loops/repetition when stuck; brittle date/regex handling; rarely uses the
llm_querysub-model call that the RLM system prompt offers. - We have not run RL on this checkpoint. No RL/GRPO numbers are claimed.
Intended use and limits
Research use: warm-starting RL, studying the RLM protocol, or as a base for further SFT. It is not a general-purpose assistant — it was trained only on long-context aggregate-reasoning tasks, and on anything else it will tend to emit REPL code. Kernels/REPL code it writes should be treated as untrusted; isolate execution.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, torch_dtype=torch.bfloat16
).to("cuda").eval()
messages = [
{"role": "user", "content": "Answer the following aggregate question.\n\nQuestion: ..."},
{"role": "user", "content": 'Example of the only accepted action format.\n\n```repl\nprint(type(context))\n```'},
{"role": "user", "content": "Turn 1/8:"},
]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(prompt, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False, pad_token_id=tok.eos_token_id)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Notes for this tokenizer: render with apply_chat_template(..., tokenize=False) and then tokenize;
generate with enable_thinking=False. generation_config stops on </assistant> (id 24) and EOS (id 2).
trust_remote_code=True is required — the custom modeling file ships in this repo.
- Downloads last month
- 416
Model tree for jucamohedano/rlm-laguna-xs2-dense-3b-v0.1
Base model
poolside/Laguna-XS.2