Qwen3-4B RLM RLVR Depth-1 LoRA Adapter

LoRA adapter from the first rank/LR ablation run.

  • Base model: Qwen/Qwen3-4B-Instruct-2507
  • Run id: rlm-rlvr-qwen3-4b-depth1-llmonly-r4-a8-lr5e-7-s150
  • Prompt variant: sanjaya_text_depth1_llm_only_v1
  • Runtime depth: depth-1 LLM-only orchestration (max_depth = 0, plain llm_query subcalls enabled)
  • LoRA rank: 4
  • LoRA alpha: 8
  • Learning rate: 5e-7
  • Training steps: 150
  • Final adapter source: run_default/broadcasts/step_150

The run_configs/ directory contains the exact trainer and orchestrator TOML files saved with the run.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr5e-7-s150-lora

Adapter
(5713)
this model

Collection including lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr5e-7-s150-lora