lsteno's picture
Fix model card formatting
a6d4b1f verified
|
Raw
History Blame Contribute Delete
710 Bytes
metadata
base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: peft
tags:
  - peft
  - lora
  - prime-rl
  - rlm-rlvr
  - qwen3

Qwen3-4B RLM RLVR Depth-1 LoRA Adapter

LoRA adapter from the first rank/LR ablation run.

  • Base model: Qwen/Qwen3-4B-Instruct-2507
  • Run id: rlm-rlvr-qwen3-4b-depth1-llmonly-r4-a8-lr5e-7-s150
  • Prompt variant: sanjaya_text_depth1_llm_only_v1
  • Runtime depth: depth-1 LLM-only orchestration (max_depth = 0, plain llm_query subcalls enabled)
  • LoRA rank: 4
  • LoRA alpha: 8
  • Learning rate: 5e-7
  • Training steps: 150
  • Final adapter source: run_default/broadcasts/step_150

The run_configs/ directory contains the exact trainer and orchestrator TOML files saved with the run.