Reinforcement Learning
PEFT
Safetensors
grpo
trl
RL Environment
openenv
p5js
generative-art
lora
sergiopaniego HF Staff commited on
Commit
511594a
·
verified ·
1 Parent(s): a6c0fd4

Drop internal run names from the card

Browse files
Files changed (1) hide show
  1. README.md +1 -2
README.md CHANGED
@@ -118,8 +118,7 @@ The base model's probe before training: reward 0.407, judge term 0.317, paint co
118
  environment and inference quota for the judge. It is not cheap.
119
 
120
  Method reproduced from [Surya Narreddi's "RL'ing Qwen to paint with
121
- code"](https://surya.website/rling-qwen-to-paint-with-code). Internally this run is `v22b`,
122
- HF job `6a95460d0718b0f6d8908805`.
123
 
124
  ## Where this comes from
125
 
 
118
  environment and inference quota for the judge. It is not cheap.
119
 
120
  Method reproduced from [Surya Narreddi's "RL'ing Qwen to paint with
121
+ code"](https://surya.website/rling-qwen-to-paint-with-code). Trained on HF Jobs (job `6a95460d0718b0f6d8908805`).
 
122
 
123
  ## Where this comes from
124