RLM-Laguna-XS.2-Dense-3B · v0.1 (protocol SFT warm start)

A ~3.0 B dense model warm-started for the Recursive Language Model (RLM) protocol: the model emits Python inside a ```repl fence, the harness executes it against an offloaded long context, and the model reads the result and decides its next action.

This checkpoint is a partial result. It reliably produces and executes valid REPL actions and attempts submissions, but it solves only a minority of the tasks we measured. See Honest status before using it.


Credits and lineage

This model would not exist without the work of others. Please credit them if you build on it.

Base model — the Poolside Laguna MoE teacher

  • poolside/Laguna-XS.2 — the 33.4 B-total / 3.0 B-active MoE (256 experts, top-8) from Poolside AI. This is the teacher model for the whole lineage.

The hackathon team — MoE-to-dense distillation and CUDA SFT

Everything between the MoE teacher and this checkpoint was produced by a hackathon team and published openly:

What that team built, which our checkpoint inherits:

  1. MoE → dense densification (K=8) with a DO-ACP expert-selection warm start rather than random init, which was the single biggest convergence lever.
  2. Reconstruction pretraining of the dense FFN against the teacher's MoE-block activations (teacher-forced, all layers in parallel).
  3. SFT on CUDA kernel data (Sakana AI-CUDA-Engineer-Archive, levels 1+2, then a follow-up SFT on levels 1+2+3) — that follow-up SFT is exactly EvanOLeary/laguna-xs2-dense-k8-cuda-sft-v2, our base model.

The RLM paradigm

  • Recursive Language Models, Alex Zhang et al. — arXiv:2512.24601, blogpost, github.com/alexzhang13/rlm. Our repository layout, environment and training stack come from that project, and the model-name convention (rlm-<base-slug>-v<major>.<minor>, cf. mit-oasys/rlm-qwen3-30b-a3b-v0.1) follows theirs.
  • prime-rl / verifiers (Prime Intellect) for the trainer and the multi-turn environment.

Datasets

  • OOLONG long-context benchmark tasks (oolongbench/oolong-synth) as the SFT/eval task family.
  • Teacher trajectories for SFT were distilled from a larger model on those tasks.

What this checkpoint adds

Starting from a CUDA-kernel SFT model, we teach the RLM action protocol. Concretely, the model learns to open a ```repl fence, write short executable Python, close the fence, and — when it believes it has the answer — set answer["ready"] = True.

Training: LoRA (r=32, α=64) on attention and FFN projections, 60 clean single-target-turn rows, resumed from an earlier adapter checkpoint. Best checkpoint by validation loss is step 50 of a 70-step run; later steps overfit.

Measured behaviour (greedy and T=0.7, real environment execution)

base (cuda-sft-v2) this checkpoint (step 50)
emits a closed ```repl fence 0 / 3 tasks 3 / 3 tasks
REPL blocks actually executed 0 3–6 per task
attempts a submission 0 / 3 2 / 3
correct answer (reward 1) 0 / 3 1 / 3

A representative first action, produced verbatim:

print(type(context))
print(len(context))
print(context[:500])

The failure mode moved from format failure to reasoning failure: the model now produces valid, executed actions, and its remaining errors are ordinary coding mistakes (for example a regex that does not match the date format) plus a tendency to loop when it cannot make progress.

Honest status

Please read this before drawing conclusions.

  • Reward is 1 on only 1 of 3 measured tasks. This is a warm start, not a finished RLM.
  • The evaluation tasks are teacher-seen. Our SFT traces were distilled over the same task pool (103 OOLONG spam tasks under our filters, of which the teacher consumed all 103). The held-out rows are held out from the SFT split, not from the teacher. These numbers are not a generalization measurement.
  • Validation loss is not a generalization number either, for the same reason. It fell 1.52 → 1.31 by step 50 and then rose to 1.62 by step 70 while training loss fell to 0.14, i.e. overfitting.
  • Known weaknesses: loops/repetition when stuck; brittle date/regex handling; rarely uses the llm_query sub-model call that the RLM system prompt offers.
  • We have not run RL on this checkpoint. No RL/GRPO numbers are claimed.

Intended use and limits

Research use: warm-starting RL, studying the RLM protocol, or as a base for further SFT. It is not a general-purpose assistant — it was trained only on long-context aggregate-reasoning tasks, and on anything else it will tend to emit REPL code. Kernels/REPL code it writes should be treated as untrusted; isolate execution.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "jucamohedano/rlm-laguna-xs2-dense-3b-v0.1"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, torch_dtype=torch.bfloat16
).to("cuda").eval()

messages = [
    {"role": "user", "content": "Answer the following aggregate question.\n\nQuestion: ..."},
    {"role": "user", "content": 'Example of the only accepted action format.\n\n```repl\nprint(type(context))\n```'},
    {"role": "user", "content": "Turn 1/8:"},
]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(prompt, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False, pad_token_id=tok.eos_token_id)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Notes for this tokenizer: render with apply_chat_template(..., tokenize=False) and then tokenize; generate with enable_thinking=False. generation_config stops on </assistant> (id 24) and EOS (id 2). trust_remote_code=True is required — the custom modeling file ships in this repo.

Downloads last month
416
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jucamohedano/rlm-laguna-xs2-dense-3b-v0.1

Paper for jucamohedano/rlm-laguna-xs2-dense-3b-v0.1