acv1229 commited on
Commit
25f6d76
·
verified ·
1 Parent(s): 374a3ef

Add README

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen2.5-Coder-7B-Instruct
3
+ tags:
4
+ - reinforcement-learning
5
+ - ppo
6
+ - lora
7
+ - code-generation
8
+ - clarification
9
+ license: apache-2.0
10
+ ---
11
+
12
+ # rl-clarify-orig-prompt-d1-0p75
13
+
14
+ LoRA fine-tune of [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct)
15
+ trained with PPO-Lagrangian constrained RL on [HumanEvalComm](https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm).
16
+
17
+ ## Training setup
18
+ - **Algorithm:** PPO with Lagrangian constraint on avg questions per episode
19
+ - **LoRA rank:** 16, alpha 32
20
+ - **Question budget (d1):** 0.75
21
+ - **Iterations:** 80
22
+ - **Checkpoint dir:** checkpoints/orig_prompt_v2/d1_0.75
23
+
24
+ ## Eval results (selected checkpoint: `iter_0039`)
25
+ - **Final eval pass@1:** 0.748 (417 problems, greedy decoding)
26
+ - **Final eval avg questions:** 0.7
27
+
28
+ ## Checkpoints
29
+ Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
30
+ with per-iteration training metrics (avg_reward, avg_questions, lambda1, lambda2, kl_per_seq).
31
+
32
+ ## Usage
33
+ ```python
34
+ from peft import PeftModel
35
+ from transformers import AutoModelForCausalLM, AutoTokenizer
36
+
37
+ base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
38
+ model = PeftModel.from_pretrained(base, "acv1229/rl-clarify-orig-prompt-d1-0p75", subfolder="iter_0039")
39
+ tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
40
+ ```