acv1229 commited on
Commit
64451c7
·
verified ·
1 Parent(s): 101b2c8

Add README

Browse files
Files changed (1) hide show
  1. README.md +38 -0
README.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen2.5-Coder-7B-Instruct
3
+ tags:
4
+ - reinforcement-learning
5
+ - ppo
6
+ - lora
7
+ - code-generation
8
+ - clarification
9
+ license: apache-2.0
10
+ ---
11
+
12
+ # rl-clarify-orig-prompt-d1-1
13
+
14
+ LoRA fine-tune of [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct)
15
+ trained with PPO-Lagrangian constrained RL on [HumanEvalComm](https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm).
16
+
17
+ PPO-Lagrangian RL fine-tune of Qwen2.5-Coder-7B-Instruct (LoRA rank 16) on HumanEvalComm with question budget d1=1. All iter_* checkpoints included. Eval checkpoint: iter_0029.
18
+
19
+ ## Training setup
20
+ - **Algorithm:** PPO with Lagrangian constraint on avg questions per episode
21
+ - **LoRA rank:** 16, alpha 32
22
+ - **Question budget (d1):** 1.0
23
+ - **Eval checkpoint:** `iter_0029`
24
+ - **Eval set:** 469 problems (full HumanEvalComm eval split, greedy decoding)
25
+
26
+ ## Checkpoints
27
+ Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
28
+ with per-iteration training metrics (avg_reward, avg_questions, lambda1).
29
+
30
+ ## Usage
31
+ ```python
32
+ from peft import PeftModel
33
+ from transformers import AutoModelForCausalLM, AutoTokenizer
34
+
35
+ base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
36
+ model = PeftModel.from_pretrained(base, "acv1229/rl-clarify-orig-prompt-d1-1", subfolder="iter_0029")
37
+ tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
38
+ ```