Add README
Browse files
README.md
CHANGED
|
@@ -27,6 +27,9 @@ PPO-Lagrangian RL fine-tune of Qwen2.5-Coder-7B-Instruct (LoRA rank 16) on Human
|
|
| 27 |
Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
|
| 28 |
with per-iteration training metrics (avg_reward, avg_questions, lambda1).
|
| 29 |
|
|
|
|
|
|
|
|
|
|
| 30 |
## Usage
|
| 31 |
```python
|
| 32 |
from peft import PeftModel
|
|
|
|
| 27 |
Each `iter_XXXX/` folder contains LoRA adapter weights and a `log.json`
|
| 28 |
with per-iteration training metrics (avg_reward, avg_questions, lambda1).
|
| 29 |
|
| 30 |
+
Training was run in two segments (job preemption on HPC); checkpoints are
|
| 31 |
+
merged from both segments in chronological order.
|
| 32 |
+
|
| 33 |
## Usage
|
| 34 |
```python
|
| 35 |
from peft import PeftModel
|