acv1229 commited on
Commit
1d692f9
·
verified ·
1 Parent(s): eeafe2b

Add README

Browse files
Files changed (1) hide show
  1. README.md +3 -0
README.md CHANGED
@@ -27,6 +27,9 @@ PPO-Lagrangian RL fine-tune of Qwen2.5-Coder-7B-Instruct (LoRA rank 16) on Human
27
  Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
28
  with per-iteration training metrics (avg_reward, avg_questions, lambda1).
29
 
 
 
 
30
  ## Usage
31
  ```python
32
  from peft import PeftModel
 
27
  Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
28
  with per-iteration training metrics (avg_reward, avg_questions, lambda1).
29
 
30
+ Training was run in two segments (job preemption on HPC); checkpoints are
31
+ merged from both segments in chronological order.
32
+
33
  ## Usage
34
  ```python
35
  from peft import PeftModel