Text Generation
PEFT
Safetensors
English
biology
single-cell
tcr
adapter
lora
chain-of-thought
reinforcement-learning
Instructions to use EthanGao123/CellHermes-CoT-RL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use EthanGao123/CellHermes-CoT-RL with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("EthanGao123/CellHermes-CoT-SFT") model = PeftModel.from_pretrained(base_model, "EthanGao123/CellHermes-CoT-RL") - Notebooks
- Google Colab
- Kaggle
Update LoRA adapter README details
Browse files
README.md
CHANGED
|
@@ -12,6 +12,10 @@ It should not be loaded directly on the original `CellHermes-v1.0` base model.
|
|
| 12 |
|
| 13 |
- Adapter files in this repository: `adapter_config.json`, `adapter_model.safetensors`
|
| 14 |
- Adapter type: LoRA
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
- Source-data variant: TCR reactivity benchmark, clone split with 25% held-out clone groups, seed 2025
|
| 16 |
- RL method: GSPO-style policy optimization with structured TCR-reactivity rewards
|
| 17 |
- Selected RL checkpoint: post-SFT RL checkpoint used for the benchmark
|
|
|
|
| 12 |
|
| 13 |
- Adapter files in this repository: `adapter_config.json`, `adapter_model.safetensors`
|
| 14 |
- Adapter type: LoRA
|
| 15 |
+
- LoRA rank: 32
|
| 16 |
+
- LoRA alpha: 64
|
| 17 |
+
- LoRA dropout: 0.0
|
| 18 |
+
- LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
|
| 19 |
- Source-data variant: TCR reactivity benchmark, clone split with 25% held-out clone groups, seed 2025
|
| 20 |
- RL method: GSPO-style policy optimization with structured TCR-reactivity rewards
|
| 21 |
- Selected RL checkpoint: post-SFT RL checkpoint used for the benchmark
|