|
Download README.md from acv1229/rl-clarify-orig-prompt-d1-0p75: direct link, hf CLI and curl.
- Browser
- Download file 1.36 kB
-
https://huggingface.co/acv1229/rl-clarify-orig-prompt-d1-0p75/resolve/main/README.md
- Command line
-
hf download hf://acv1229/rl-clarify-orig-prompt-d1-0p75/README.md
-
curl -L -o README.md https://huggingface.co/acv1229/rl-clarify-orig-prompt-d1-0p75/resolve/main/README.md
1.36 kB
| base_model: Qwen/Qwen2.5-Coder-7B-Instruct | |
| tags: | |
| - reinforcement-learning | |
| - ppo | |
| - lora | |
| - code-generation | |
| - clarification | |
| license: apache-2.0 | |
| # rl-clarify-orig-prompt-d1-0p75 | |
| LoRA fine-tune of [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct) | |
| trained with PPO-Lagrangian constrained RL on [HumanEvalComm](https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm). | |
| ## Training setup | |
| - **Algorithm:** PPO with Lagrangian constraint on avg questions per episode | |
| - **LoRA rank:** 16, alpha 32 | |
| - **Question budget (d1):** 0.75 | |
| - **Iterations:** 80 | |
| - **Checkpoint dir:** checkpoints/orig_prompt_v2/d1_0.75 | |
| ## Eval results (selected checkpoint: `iter_0039`) | |
| - **Final eval pass@1:** 0.748 (417 problems, greedy decoding) | |
| - **Final eval avg questions:** 0.7 | |
| ## Checkpoints | |
| Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json` | |
| with per-iteration training metrics (avg_reward, avg_questions, lambda1, lambda2, kl_per_seq). | |
| ## Usage | |
| ```python | |
| from peft import PeftModel | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct") | |
| model = PeftModel.from_pretrained(base, "acv1229/rl-clarify-orig-prompt-d1-0p75", subfolder="iter_0039") | |
| tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct") | |
| ``` | |