|
Download README.md from acv1229/rl-clarify-orig-prompt-d1-1p5: direct link, hf CLI and curl.
- Browser
- Download file 1.35 kB
-
https://huggingface.co/acv1229/rl-clarify-orig-prompt-d1-1p5/resolve/main/README.md
- Command line
-
hf download hf://acv1229/rl-clarify-orig-prompt-d1-1p5/README.md
-
curl -L -o README.md https://huggingface.co/acv1229/rl-clarify-orig-prompt-d1-1p5/resolve/main/README.md
1.35 kB
metadata
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
tags:
- reinforcement-learning
- ppo
- lora
- code-generation
- clarification
license: apache-2.0
rl-clarify-orig-prompt-d1-1p5
LoRA fine-tune of Qwen2.5-Coder-7B-Instruct trained with PPO-Lagrangian constrained RL on HumanEvalComm.
Training setup
- Algorithm: PPO with Lagrangian constraint on avg questions per episode
- LoRA rank: 16, alpha 32
- Question budget (d1): 1.5
- Iterations: 80
- Checkpoint dir: checkpoints/orig_prompt/d1_1.5
Eval results (selected checkpoint: iter_0049)
- Final eval pass@1: 0.759 (417 problems, greedy decoding)
- Final eval avg questions: 0.8
Checkpoints
Each iter_XXXX/ folder contains LoRA adapter weights and a log.json
with per-iteration training metrics (avg_reward, avg_questions, lambda1, lambda2, kl_per_seq).
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
model = PeftModel.from_pretrained(base, "acv1229/rl-clarify-orig-prompt-d1-1p5", subfolder="iter_0049")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")