thealper2 commited on
Commit
1eb78b1
·
verified ·
1 Parent(s): e29e644

Upload fine-tuned T5-small commit message generator

Browse files
README.md ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google-t5/t5-small
4
+ tags:
5
+ - commit-message-generation
6
+ - text2text-generation
7
+ - summarization
8
+ - code
9
+ datasets:
10
+ - Maxscha/commitbench
11
+ language:
12
+ - en
13
+ library_name: transformers
14
+ pipeline_tag: text2text-generation
15
+ metrics:
16
+ - rouge
17
+ - bleu
18
+ ---
19
+
20
+ # thealper2/t5-small-commitbench
21
+
22
+ `google-t5/t5-small` fine-tuned on [Maxscha/commitbench](https://huggingface.co/datasets/Maxscha/commitbench) for
23
+ commit message generation: given a git diff, generate the commit message describing it.
24
+
25
+ ## Task format
26
+
27
+ Text-to-text. The input is a task prefix followed by the raw git diff, the target is the
28
+ commit message.
29
+
30
+ ```
31
+ generate commit message: <git diff>
32
+ ```
33
+
34
+ ## Training data
35
+
36
+ [Maxscha/commitbench](https://huggingface.co/datasets/Maxscha/commitbench) official splits, used unchanged:
37
+
38
+ | Split | Examples in split | Examples used |
39
+ |---|---|---|
40
+ | train | 1,165,213 | 500,000 |
41
+ | validation | 249,689 | 2,000 |
42
+ | test | 249,688 | not used for training |
43
+
44
+ Languages covered by the dataset: Python, JavaScript, PHP, Ruby, Java, Go.
45
+
46
+ ## Training configuration
47
+
48
+ | Setting | Value |
49
+ |---|---|
50
+ | Base model | `google-t5/t5-small` |
51
+ | Parameters | 60.5M |
52
+ | Max source length | 512 tokens |
53
+ | Max target length | 64 tokens |
54
+ | Per-device batch size | 32 |
55
+ | Gradient accumulation | 1 |
56
+ | Effective batch size | 32 |
57
+ | Learning rate | 3e-05 |
58
+ | LR schedule | linear |
59
+ | Warmup ratio | 0.05 |
60
+ | Weight decay | 0.01 |
61
+ | Epochs | 2.0 |
62
+ | Label smoothing | 0.0 |
63
+ | Gradient clipping | 1.0 |
64
+ | Mixed precision | bf16 |
65
+ | Seed | 42 |
66
+ | Optimizer | AdamW |
67
+ | Training time | 1.219 h |
68
+ | Hardware | NVIDIA GeForce RTX 5060 Ti (15.9 GB) |
69
+
70
+ Truncation at these limits (measured on a 50k sample with the T5 tokenizer):
71
+ - 0.7% of the diffs exceed 512 source tokens.
72
+ - 4.47% of the commit messages exceed 64 target tokens.
73
+
74
+ ## Results
75
+
76
+ - Final training loss: **3.5762**
77
+ - Best validation loss: **3.2414**
78
+
79
+ Test split (20,000 examples), beam search with `num_beams=4`:
80
+
81
+ | Metric | Value |
82
+ |---|---|
83
+ | rouge1 | 19.31 |
84
+ | rouge2 | 4.668 |
85
+ | rougeL | 17.42 |
86
+ | rougeLsum | 17.42 |
87
+ | bleu | 2.148 |
88
+ | exact_match | 0.04 |
89
+ | gen_len_words_mean | 5.005 |
90
+ | ref_len_words_mean | 11.27 |
91
+
92
+ Per programming language:
93
+
94
+ | Language | n | ROUGE-1 | ROUGE-2 | ROUGE-L | BLEU | Exact match |
95
+ |---|---|---|---|---|---|---|
96
+ | Python | 5,722 | 21.20 | 6.05 | 19.29 | 2.73 | 0.04 |
97
+ | JavaScript | 4,468 | 18.86 | 4.05 | 17.07 | 2.01 | 0.02 |
98
+ | PHP | 3,489 | 17.04 | 3.46 | 15.29 | 1.64 | 0.09 |
99
+ | Ruby | 2,808 | 22.08 | 5.79 | 19.65 | 2.44 | 0.04 |
100
+ | Java | 1,799 | 15.19 | 2.61 | 13.58 | 1.02 | 0.06 |
101
+ | Go | 1,714 | 18.65 | 4.45 | 16.75 | 2.11 | 0.00 |
102
+
103
+ ROUGE and BLEU are lexical-overlap metrics. They do not fully capture whether a commit
104
+ message describes a change correctly, and generic messages can score well.
105
+
106
+ ## Usage
107
+
108
+ ```python
109
+ from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
110
+
111
+ model_id = "thealper2/t5-small-commitbench"
112
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
113
+ model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
114
+
115
+ diff = open("change.patch").read()
116
+ inputs = tokenizer(
117
+ "generate commit message: " + diff,
118
+ max_length=512,
119
+ truncation=True,
120
+ return_tensors="pt",
121
+ )
122
+ output = model.generate(
123
+ **inputs,
124
+ num_beams=4,
125
+ max_new_tokens=64,
126
+ length_penalty=1.0,
127
+ no_repeat_ngram_size=3,
128
+ )
129
+ print(tokenizer.decode(output[0], skip_special_tokens=True))
130
+ ```
131
+
132
+ Default generation settings: `num_beams=4`, `max_new_tokens=64`,
133
+ `min_new_tokens=0`, `length_penalty=1.0`,
134
+ `no_repeat_ngram_size=3`, `do_sample=False` (deterministic).
135
+
136
+ ## Limitations
137
+
138
+ - CommitBench replaces identifying literals with placeholder tokens: every diff contains
139
+ `<HASH>` instead of commit hashes, and 26.5% of the reference messages contain `<I>` (numbers),
140
+ `<URL>` or `<EMAIL>`. The model therefore also generates these tokens, e.g. `Bumped version to <I>`.
141
+ - The T5 sentencepiece vocabulary does not cover every character used in source code (curly braces, backslashes, angle brackets), so about 2.35% of the input tokens become `<unk>`. This limits how precisely the model can read a diff.
142
+ - Diffs longer than 512 tokens are truncated; the tail of the change is
143
+ not visible to the model.
144
+ - CommitBench splits are random over commits, not over repositories: 98.6% of the test examples
145
+ come from repositories that also appear in the training split. No `(diff, message)` pair is
146
+ shared across splits, but the reported scores partly reflect familiarity with a project's
147
+ commit style rather than generalization to unseen code.
148
+ - The dataset is English-only and covers six languages; behaviour on other languages or
149
+ on very large multi-file changes is untested.
150
+ - CommitBench is released under CC BY-NC 4.0, which restricts commercial use of the data.
151
+
152
+ ## Reproducibility
153
+
154
+ - python: `3.12.3`
155
+ - torch: `2.11.0+cu128`
156
+ - transformers: `5.17.0`
157
+ - datasets: `4.3.0`
158
+ - tokenizers: `0.23.2`
159
+ - seed: `42`
config.json ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "T5ForConditionalGeneration"
4
+ ],
5
+ "classifier_dropout": 0.0,
6
+ "d_ff": 2048,
7
+ "d_kv": 64,
8
+ "d_model": 512,
9
+ "decoder_start_token_id": 0,
10
+ "dense_act_fn": "relu",
11
+ "dropout_rate": 0.1,
12
+ "dtype": "float32",
13
+ "eos_token_id": 1,
14
+ "feed_forward_proj": "relu",
15
+ "initializer_factor": 1.0,
16
+ "is_decoder": false,
17
+ "is_encoder_decoder": true,
18
+ "is_gated_act": false,
19
+ "layer_norm_epsilon": 1e-06,
20
+ "model_type": "t5",
21
+ "n_positions": 512,
22
+ "num_decoder_layers": 6,
23
+ "num_heads": 8,
24
+ "num_layers": 6,
25
+ "output_past": true,
26
+ "pad_token_id": 0,
27
+ "relative_attention_max_distance": 128,
28
+ "relative_attention_num_buckets": 32,
29
+ "scale_decoder_outputs": true,
30
+ "task_specific_params": {
31
+ "summarization": {
32
+ "early_stopping": true,
33
+ "length_penalty": 2.0,
34
+ "max_length": 200,
35
+ "min_length": 30,
36
+ "no_repeat_ngram_size": 3,
37
+ "num_beams": 4,
38
+ "prefix": "summarize: "
39
+ },
40
+ "translation_en_to_de": {
41
+ "early_stopping": true,
42
+ "max_length": 300,
43
+ "num_beams": 4,
44
+ "prefix": "translate English to German: "
45
+ },
46
+ "translation_en_to_fr": {
47
+ "early_stopping": true,
48
+ "max_length": 300,
49
+ "num_beams": 4,
50
+ "prefix": "translate English to French: "
51
+ },
52
+ "translation_en_to_ro": {
53
+ "early_stopping": true,
54
+ "max_length": 300,
55
+ "num_beams": 4,
56
+ "prefix": "translate English to Romanian: "
57
+ }
58
+ },
59
+ "tie_word_embeddings": true,
60
+ "transformers_version": "5.17.0",
61
+ "use_cache": true,
62
+ "vocab_size": 32128
63
+ }
generation_config.json ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "assistant_confidence_threshold": 0.4,
3
+ "assistant_lookbehind": 10,
4
+ "decoder_start_token_id": 0,
5
+ "diversity_penalty": 0.0,
6
+ "do_sample": false,
7
+ "early_stopping": true,
8
+ "encoder_no_repeat_ngram_size": 0,
9
+ "encoder_repetition_penalty": 1.0,
10
+ "eos_token_id": [
11
+ 1
12
+ ],
13
+ "epsilon_cutoff": 0.0,
14
+ "eta_cutoff": 0.0,
15
+ "length_penalty": 1.0,
16
+ "max_length": 20,
17
+ "max_new_tokens": 64,
18
+ "min_length": 0,
19
+ "min_new_tokens": 0,
20
+ "no_repeat_ngram_size": 3,
21
+ "num_assistant_tokens": 20,
22
+ "num_assistant_tokens_schedule": "constant",
23
+ "num_beam_groups": 1,
24
+ "num_beams": 4,
25
+ "num_return_sequences": 1,
26
+ "output_scores": false,
27
+ "pad_token_id": 0,
28
+ "remove_invalid_values": false,
29
+ "repetition_penalty": 1.0,
30
+ "return_dict_in_generate": false,
31
+ "target_lookbehind": 10,
32
+ "temperature": 1.0,
33
+ "top_k": 50,
34
+ "top_p": 1.0,
35
+ "transformers_version": "5.17.0",
36
+ "typical_p": 1.0,
37
+ "use_cache": true
38
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0d619724d679d2fb7bd6e597b09d4af441c14e37ded9e918868ba63f58b40de4
3
+ size 242041896
run_config.json ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "data": {
3
+ "dataset_name": "Maxscha/commitbench",
4
+ "cache_dir": null,
5
+ "prefix": "generate commit message: ",
6
+ "target_mode": "full",
7
+ "max_source_length": 512,
8
+ "max_target_length": 64,
9
+ "min_diff_chars": 1,
10
+ "min_message_chars": 1,
11
+ "num_proc": 8,
12
+ "max_train_samples": 500000,
13
+ "max_eval_samples": 2000,
14
+ "max_predict_samples": null
15
+ },
16
+ "train": {
17
+ "model_name": "google-t5/t5-small",
18
+ "output_dir": "outputs/t5-small-commitbench",
19
+ "seed": 42,
20
+ "deterministic": false,
21
+ "learning_rate": 3e-05,
22
+ "num_train_epochs": 2.0,
23
+ "per_device_train_batch_size": 32,
24
+ "per_device_eval_batch_size": 64,
25
+ "gradient_accumulation_steps": 1,
26
+ "weight_decay": 0.01,
27
+ "warmup_ratio": 0.05,
28
+ "label_smoothing_factor": 0.0,
29
+ "max_grad_norm": 1.0,
30
+ "lr_scheduler_type": "linear",
31
+ "eval_strategy": "steps",
32
+ "save_strategy": "steps",
33
+ "eval_steps": 3000,
34
+ "save_steps": 3000,
35
+ "logging_steps": 250,
36
+ "save_total_limit": 2,
37
+ "metric_for_best_model": "eval_loss",
38
+ "greater_is_better": false,
39
+ "load_best_model_at_end": true,
40
+ "gradient_checkpointing": false,
41
+ "torch_compile": true,
42
+ "torch_compile_mode": "",
43
+ "group_by_length": true,
44
+ "dataloader_num_workers": 4,
45
+ "precision": "auto",
46
+ "resume_from_checkpoint": null,
47
+ "max_steps": -1
48
+ },
49
+ "generation": {
50
+ "num_beams": 4,
51
+ "max_new_tokens": 64,
52
+ "min_new_tokens": 0,
53
+ "length_penalty": 1.0,
54
+ "no_repeat_ngram_size": 3,
55
+ "early_stopping": true,
56
+ "do_sample": false,
57
+ "eval_num_beams": 1
58
+ }
59
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": true,
4
+ "eos_token": "</s>",
5
+ "extra_ids": 100,
6
+ "extra_special_tokens": [
7
+ "<extra_id_0>",
8
+ "<extra_id_1>",
9
+ "<extra_id_2>",
10
+ "<extra_id_3>",
11
+ "<extra_id_4>",
12
+ "<extra_id_5>",
13
+ "<extra_id_6>",
14
+ "<extra_id_7>",
15
+ "<extra_id_8>",
16
+ "<extra_id_9>",
17
+ "<extra_id_10>",
18
+ "<extra_id_11>",
19
+ "<extra_id_12>",
20
+ "<extra_id_13>",
21
+ "<extra_id_14>",
22
+ "<extra_id_15>",
23
+ "<extra_id_16>",
24
+ "<extra_id_17>",
25
+ "<extra_id_18>",
26
+ "<extra_id_19>",
27
+ "<extra_id_20>",
28
+ "<extra_id_21>",
29
+ "<extra_id_22>",
30
+ "<extra_id_23>",
31
+ "<extra_id_24>",
32
+ "<extra_id_25>",
33
+ "<extra_id_26>",
34
+ "<extra_id_27>",
35
+ "<extra_id_28>",
36
+ "<extra_id_29>",
37
+ "<extra_id_30>",
38
+ "<extra_id_31>",
39
+ "<extra_id_32>",
40
+ "<extra_id_33>",
41
+ "<extra_id_34>",
42
+ "<extra_id_35>",
43
+ "<extra_id_36>",
44
+ "<extra_id_37>",
45
+ "<extra_id_38>",
46
+ "<extra_id_39>",
47
+ "<extra_id_40>",
48
+ "<extra_id_41>",
49
+ "<extra_id_42>",
50
+ "<extra_id_43>",
51
+ "<extra_id_44>",
52
+ "<extra_id_45>",
53
+ "<extra_id_46>",
54
+ "<extra_id_47>",
55
+ "<extra_id_48>",
56
+ "<extra_id_49>",
57
+ "<extra_id_50>",
58
+ "<extra_id_51>",
59
+ "<extra_id_52>",
60
+ "<extra_id_53>",
61
+ "<extra_id_54>",
62
+ "<extra_id_55>",
63
+ "<extra_id_56>",
64
+ "<extra_id_57>",
65
+ "<extra_id_58>",
66
+ "<extra_id_59>",
67
+ "<extra_id_60>",
68
+ "<extra_id_61>",
69
+ "<extra_id_62>",
70
+ "<extra_id_63>",
71
+ "<extra_id_64>",
72
+ "<extra_id_65>",
73
+ "<extra_id_66>",
74
+ "<extra_id_67>",
75
+ "<extra_id_68>",
76
+ "<extra_id_69>",
77
+ "<extra_id_70>",
78
+ "<extra_id_71>",
79
+ "<extra_id_72>",
80
+ "<extra_id_73>",
81
+ "<extra_id_74>",
82
+ "<extra_id_75>",
83
+ "<extra_id_76>",
84
+ "<extra_id_77>",
85
+ "<extra_id_78>",
86
+ "<extra_id_79>",
87
+ "<extra_id_80>",
88
+ "<extra_id_81>",
89
+ "<extra_id_82>",
90
+ "<extra_id_83>",
91
+ "<extra_id_84>",
92
+ "<extra_id_85>",
93
+ "<extra_id_86>",
94
+ "<extra_id_87>",
95
+ "<extra_id_88>",
96
+ "<extra_id_89>",
97
+ "<extra_id_90>",
98
+ "<extra_id_91>",
99
+ "<extra_id_92>",
100
+ "<extra_id_93>",
101
+ "<extra_id_94>",
102
+ "<extra_id_95>",
103
+ "<extra_id_96>",
104
+ "<extra_id_97>",
105
+ "<extra_id_98>",
106
+ "<extra_id_99>"
107
+ ],
108
+ "is_local": false,
109
+ "local_files_only": false,
110
+ "model_max_length": 512,
111
+ "pad_token": "<pad>",
112
+ "tokenizer_class": "T5Tokenizer",
113
+ "unk_token": "<unk>"
114
+ }