Instructions to use FineEnvs/watercolour-grpo-judge-led with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use FineEnvs/watercolour-grpo-judge-led with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-35B-A3B") model = PeftModel.from_pretrained(base_model, "FineEnvs/watercolour-grpo-judge-led") - Notebooks
- Google Colab
- Kaggle
Drop internal run names from the card
Browse files
README.md
CHANGED
|
@@ -118,8 +118,7 @@ The base model's probe before training: reward 0.407, judge term 0.317, paint co
|
|
| 118 |
environment and inference quota for the judge. It is not cheap.
|
| 119 |
|
| 120 |
Method reproduced from [Surya Narreddi's "RL'ing Qwen to paint with
|
| 121 |
-
code"](https://surya.website/rling-qwen-to-paint-with-code).
|
| 122 |
-
HF job `6a95460d0718b0f6d8908805`.
|
| 123 |
|
| 124 |
## Where this comes from
|
| 125 |
|
|
|
|
| 118 |
environment and inference quota for the judge. It is not cheap.
|
| 119 |
|
| 120 |
Method reproduced from [Surya Narreddi's "RL'ing Qwen to paint with
|
| 121 |
+
code"](https://surya.website/rling-qwen-to-paint-with-code). Trained on HF Jobs (job `6a95460d0718b0f6d8908805`).
|
|
|
|
| 122 |
|
| 123 |
## Where this comes from
|
| 124 |
|