Vidushee commited on
Commit
b8c00db
·
verified ·
1 Parent(s): b055905

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +101 -0
README.md ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3-32B
4
+ datasets:
5
+ - Vidushee/BT_Preference_Dataset
6
+ pipeline_tag: text-classification
7
+ tags:
8
+ - reward-model
9
+ - bradley-terry
10
+ - rlhf
11
+ ---
12
+
13
+ # Qwen3-32B Bradley-Terry Reward Model
14
+
15
+ A Bradley-Terry reward model fine-tuned from [Qwen/Qwen3-32B](https://huggingface.co/Qwen/Qwen3-32B) for scoring question quality about research papers.
16
+
17
+ ## Training Details
18
+
19
+ - **Base model**: Qwen/Qwen3-32B
20
+ - **Dataset**: [Vidushee/BT_Preference_Dataset](https://huggingface.co/datasets/Vidushee/BT_Preference_Dataset) (28,049 train, 3,090 test pairs)
21
+ - **Training**: 1 epoch, 400/877 steps, batch size 8 (1 per device x 8 gradient accumulation x 4 GPUs)
22
+ - **Hardware**: 4x NVIDIA H100 80GB GPUs (single node)
23
+ - **Framework**: HuggingFace Trainer + DeepSpeed ZeRO-3 + CPU optimizer offload
24
+ - **Learning rate**: 1e-6 with cosine schedule and 3% warmup
25
+ - **Max sequence length**: 12,288 tokens
26
+ - **Eval accuracy**: 90.9% pairwise accuracy at step 400
27
+ - **Training approach**: Adapted from [RLHFlow/RLHF-Reward-Modeling](https://github.com/RLHFlow/RLHF-Reward-Modeling/tree/main/bradley-terry-rm)
28
+
29
+ ## Usage
30
+
31
+ ```python
32
+ import re
33
+ import torch
34
+ from transformers import AutoTokenizer, pipeline
35
+
36
+ model_path = "Vidushee/Qwen3-32B-BT-RewardModel"
37
+ tokenizer = AutoTokenizer.from_pretrained(model_path)
38
+
39
+ rm_pipe = pipeline(
40
+ "sentiment-analysis",
41
+ model=model_path,
42
+ device=0,
43
+ tokenizer=tokenizer,
44
+ model_kwargs={"torch_dtype": torch.bfloat16, "attn_implementation": "flash_attention_2"},
45
+ truncation=True,
46
+ max_length=12288,
47
+ )
48
+
49
+ pipe_kwargs = {
50
+ "return_all_scores": True,
51
+ "function_to_apply": "none",
52
+ "batch_size": 1,
53
+ }
54
+
55
+ # Format your conversation
56
+ chat = [
57
+ {"role": "user", "content": "Your paper context here"},
58
+ {"role": "assistant", "content": "Question to score"},
59
+ ]
60
+
61
+ text = tokenizer.apply_chat_template(
62
+ chat, tokenize=False, add_generation_prompt=False, enable_thinking=False
63
+ )
64
+ # Strip empty think blocks that Qwen3 inserts even with enable_thinking=False
65
+ text = re.sub(r"<think>\s*</think>\s*", "", text)
66
+ # Strip trailing newline so reward pools from <|im_end|>
67
+ text = text.rstrip("\n")
68
+
69
+ outputs = rm_pipe([text], **pipe_kwargs)
70
+ reward = outputs[0][0]["score"]
71
+ print(f"Reward: {reward}")
72
+ ```
73
+
74
+ ## Comparing Two Responses
75
+
76
+ ```python
77
+ # Score chosen vs rejected responses
78
+ chosen_chat = [
79
+ {"role": "user", "content": "Paper context..."},
80
+ {"role": "assistant", "content": "Good question about the paper"},
81
+ ]
82
+ rejected_chat = [
83
+ {"role": "user", "content": "Paper context..."},
84
+ {"role": "assistant", "content": "Bad question about the paper"},
85
+ ]
86
+
87
+ def format_text(messages):
88
+ text = tokenizer.apply_chat_template(
89
+ messages, tokenize=False, add_generation_prompt=False, enable_thinking=False
90
+ )
91
+ text = re.sub(r"<think>\s*</think>\s*", "", text)
92
+ return text.rstrip("\n")
93
+
94
+ outputs = rm_pipe([format_text(chosen_chat), format_text(rejected_chat)], **pipe_kwargs)
95
+ chosen_reward = outputs[0][0]["score"]
96
+ rejected_reward = outputs[1][0]["score"]
97
+
98
+ print(f"Chosen reward: {chosen_reward:.4f}")
99
+ print(f"Rejected reward: {rejected_reward:.4f}")
100
+ print(f"Chosen is better: {chosen_reward > rejected_reward}")
101
+ ```