suehuynh commited on
Commit
78db8a3
·
verified ·
1 Parent(s): 9bdef50

Delete adaption_mixtral_8x7b_instruct_marketing_tasks

Browse files
adaption_mixtral_8x7b_instruct_marketing_tasks/README.md DELETED
@@ -1,202 +0,0 @@
1
- ---
2
- base_model: mistralai/Mixtral-8x7B-Instruct-v0.1
3
- library_name: peft
4
- ---
5
-
6
- # Model Card for Model ID
7
-
8
- <!-- Provide a quick summary of what the model is/does. -->
9
-
10
-
11
-
12
- ## Model Details
13
-
14
- ### Model Description
15
-
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
-
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
-
28
- ### Model Sources [optional]
29
-
30
- <!-- Provide the basic links for the model. -->
31
-
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
-
36
- ## Uses
37
-
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
- ### Direct Use
41
-
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
-
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
102
-
103
- ## Evaluation
104
-
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
- ### Testing Data, Factors & Metrics
108
-
109
- #### Testing Data
110
-
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
-
121
- #### Metrics
122
-
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
126
-
127
- ### Results
128
-
129
- [More Information Needed]
130
-
131
- #### Summary
132
-
133
-
134
-
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
-
139
- [More Information Needed]
140
-
141
- ## Environmental Impact
142
-
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
-
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
-
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
-
153
- ## Technical Specifications [optional]
154
-
155
- ### Model Architecture and Objective
156
-
157
- [More Information Needed]
158
-
159
- ### Compute Infrastructure
160
-
161
- [More Information Needed]
162
-
163
- #### Hardware
164
-
165
- [More Information Needed]
166
-
167
- #### Software
168
-
169
- [More Information Needed]
170
-
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
-
175
- **BibTeX:**
176
-
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
200
- ### Framework versions
201
-
202
- - PEFT 0.15.1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/adapter_config.json DELETED
@@ -1,36 +0,0 @@
1
- {
2
- "alpha_pattern": {},
3
- "auto_mapping": null,
4
- "base_model_name_or_path": "mistralai/Mixtral-8x7B-Instruct-v0.1",
5
- "bias": "none",
6
- "corda_config": null,
7
- "eva_config": null,
8
- "exclude_modules": [],
9
- "fan_in_fan_out": false,
10
- "inference_mode": true,
11
- "init_lora_weights": true,
12
- "layer_replication": null,
13
- "layers_pattern": null,
14
- "layers_to_transform": null,
15
- "loftq_config": {},
16
- "lora_alpha": 32,
17
- "lora_bias": false,
18
- "lora_dropout": 0.0,
19
- "megatron_config": null,
20
- "megatron_core": "megatron.core",
21
- "modules_to_save": null,
22
- "peft_type": "LORA",
23
- "r": 16,
24
- "rank_pattern": {},
25
- "revision": null,
26
- "target_modules": [
27
- "v_proj",
28
- "q_proj",
29
- "k_proj",
30
- "o_proj"
31
- ],
32
- "task_type": "CAUSAL_LM",
33
- "trainable_token_indices": null,
34
- "use_dora": false,
35
- "use_rslora": false
36
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/adapter_model.safetensors DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:e9b5fb63422cf56622ffb54c8d6b23163eb5a416c3588d7b0f0abdf6d66777af
3
- size 54560368
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/chat_template.jinja DELETED
@@ -1,24 +0,0 @@
1
- {%- if messages[0]['role'] == 'system' %}
2
- {%- set system_message = messages[0]['content'] %}
3
- {%- set loop_messages = messages[1:] %}
4
- {%- else %}
5
- {%- set loop_messages = messages %}
6
- {%- endif %}
7
-
8
- {{- bos_token }}
9
- {%- for message in loop_messages %}
10
- {%- if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}
11
- {{- raise_exception('After the optional system message, conversation roles must alternate user/assistant/user/assistant/...') }}
12
- {%- endif %}
13
- {%- if message['role'] == 'user' %}
14
- {%- if loop.first and system_message is defined %}
15
- {{- ' [INST] ' + system_message + '\n\n' + message['content'] + ' [/INST]' }}
16
- {%- else %}
17
- {{- ' [INST] ' + message['content'] + ' [/INST]' }}
18
- {%- endif %}
19
- {%- elif message['role'] == 'assistant' %}
20
- {{- ' ' + message['content'] + eos_token}}
21
- {%- else %}
22
- {{- raise_exception('Only user and assistant roles are supported, with the exception of an initial optional system message!') }}
23
- {%- endif %}
24
- {%- endfor %}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/config.json DELETED
@@ -1,36 +0,0 @@
1
- {
2
- "architectures": [
3
- "MixtralForCausalLM"
4
- ],
5
- "attention_dropout": 0.0,
6
- "bos_token_id": 1,
7
- "dtype": "bfloat16",
8
- "eos_token_id": 2,
9
- "head_dim": null,
10
- "hidden_act": "silu",
11
- "hidden_size": 4096,
12
- "initializer_range": 0.02,
13
- "intermediate_size": 14336,
14
- "max_position_embeddings": 32768,
15
- "model_type": "mixtral",
16
- "num_attention_heads": 32,
17
- "num_experts_per_tok": 2,
18
- "num_hidden_layers": 32,
19
- "num_key_value_heads": 8,
20
- "num_local_experts": 8,
21
- "output_router_logits": false,
22
- "pad_token_id": 2,
23
- "rms_norm_eps": 1e-05,
24
- "rope_parameters": {
25
- "rope_theta": 1000000.0,
26
- "rope_type": "default"
27
- },
28
- "router_aux_loss_coef": 0.02,
29
- "router_jitter_noise": 0.0,
30
- "sliding_window": null,
31
- "tie_word_embeddings": false,
32
- "transformers_version": "5.10.1",
33
- "use_cache": false,
34
- "vocab_size": 32000,
35
- "torch_dtype": "bfloat16"
36
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/tokenizer.json DELETED
The diff for this file is too large to render. See raw diff
 
adaption_mixtral_8x7b_instruct_marketing_tasks/tokenizer_config.json DELETED
@@ -1,19 +0,0 @@
1
- {
2
- "add_prefix_space": null,
3
- "backend": "tokenizers",
4
- "bos_token": "<s>",
5
- "clean_up_tokenization_spaces": false,
6
- "eos_token": "</s>",
7
- "extra_special_tokens": [],
8
- "is_local": false,
9
- "legacy": false,
10
- "local_files_only": true,
11
- "model_max_length": 32768,
12
- "pad_token": "</s>",
13
- "padding_side": "right",
14
- "sp_model_kwargs": {},
15
- "spaces_between_special_tokens": false,
16
- "tokenizer_class": "TokenizersBackend",
17
- "unk_token": "<unk>",
18
- "use_default_system_prompt": false
19
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adaption_mixtral_8x7b_instruct_marketing_tasks/trainer_state.json DELETED
@@ -1,368 +0,0 @@
1
- {
2
- "best_global_step": null,
3
- "best_metric": null,
4
- "best_model_checkpoint": null,
5
- "epoch": 1.0,
6
- "eval_steps": 8,
7
- "global_step": 42,
8
- "is_hyper_param_search": false,
9
- "is_local_process_zero": true,
10
- "is_world_process_zero": true,
11
- "log_history": [
12
- {
13
- "epoch": 0.023809523809523808,
14
- "grad_norm": 0.5433083772659302,
15
- "learning_rate": 0.0,
16
- "loss": 1.4814453125,
17
- "step": 1
18
- },
19
- {
20
- "epoch": 0.047619047619047616,
21
- "grad_norm": 0.42392125725746155,
22
- "learning_rate": 2e-05,
23
- "loss": 1.35400390625,
24
- "step": 2
25
- },
26
- {
27
- "epoch": 0.07142857142857142,
28
- "grad_norm": 0.47326141595840454,
29
- "learning_rate": 4e-05,
30
- "loss": 1.4248046875,
31
- "step": 3
32
- },
33
- {
34
- "epoch": 0.09523809523809523,
35
- "grad_norm": 0.535437285900116,
36
- "learning_rate": 6e-05,
37
- "loss": 1.5166015625,
38
- "step": 4
39
- },
40
- {
41
- "epoch": 0.11904761904761904,
42
- "grad_norm": 0.601170003414154,
43
- "learning_rate": 8e-05,
44
- "loss": 1.490234375,
45
- "step": 5
46
- },
47
- {
48
- "epoch": 0.14285714285714285,
49
- "grad_norm": 0.4886681139469147,
50
- "learning_rate": 0.0001,
51
- "loss": 1.40625,
52
- "step": 6
53
- },
54
- {
55
- "epoch": 0.16666666666666666,
56
- "grad_norm": 0.40877094864845276,
57
- "learning_rate": 9.983788698441369e-05,
58
- "loss": 1.2783203125,
59
- "step": 7
60
- },
61
- {
62
- "epoch": 0.19047619047619047,
63
- "grad_norm": 0.3925204277038574,
64
- "learning_rate": 9.935271596564688e-05,
65
- "loss": 1.15478515625,
66
- "step": 8
67
- },
68
- {
69
- "epoch": 0.21428571428571427,
70
- "grad_norm": 0.2866791784763336,
71
- "learning_rate": 9.854798261200746e-05,
72
- "loss": 1.220703125,
73
- "step": 9
74
- },
75
- {
76
- "epoch": 0.23809523809523808,
77
- "grad_norm": 0.11159192025661469,
78
- "learning_rate": 9.74294850457488e-05,
79
- "loss": 1.1728515625,
80
- "step": 10
81
- },
82
- {
83
- "epoch": 0.23809523809523808,
84
- "eval_loss": 1.0947265625,
85
- "eval_runtime": 2.2064,
86
- "eval_samples_per_second": 2.719,
87
- "eval_steps_per_second": 0.453,
88
- "step": 10
89
- },
90
- {
91
- "epoch": 0.2619047619047619,
92
- "grad_norm": 0.10726847499608994,
93
- "learning_rate": 9.600528206746612e-05,
94
- "loss": 0.996826171875,
95
- "step": 11
96
- },
97
- {
98
- "epoch": 0.2857142857142857,
99
- "grad_norm": 0.08991753309965134,
100
- "learning_rate": 9.428563509225347e-05,
101
- "loss": 1.11279296875,
102
- "step": 12
103
- },
104
- {
105
- "epoch": 0.30952380952380953,
106
- "grad_norm": 0.0953637957572937,
107
- "learning_rate": 9.22829342159729e-05,
108
- "loss": 1.09423828125,
109
- "step": 13
110
- },
111
- {
112
- "epoch": 0.3333333333333333,
113
- "grad_norm": 0.09409310668706894,
114
- "learning_rate": 9.001160894432978e-05,
115
- "loss": 1.06494140625,
116
- "step": 14
117
- },
118
- {
119
- "epoch": 0.35714285714285715,
120
- "grad_norm": 0.10521717369556427,
121
- "learning_rate": 8.74880242279536e-05,
122
- "loss": 1.1982421875,
123
- "step": 15
124
- },
125
- {
126
- "epoch": 0.38095238095238093,
127
- "grad_norm": 0.0813065692782402,
128
- "learning_rate": 8.473036255255366e-05,
129
- "loss": 0.9486083984375,
130
- "step": 16
131
- },
132
- {
133
- "epoch": 0.40476190476190477,
134
- "grad_norm": 0.09203798323869705,
135
- "learning_rate": 8.175849293369291e-05,
136
- "loss": 1.0810546875,
137
- "step": 17
138
- },
139
- {
140
- "epoch": 0.42857142857142855,
141
- "grad_norm": 0.08192384988069534,
142
- "learning_rate": 7.859382776007543e-05,
143
- "loss": 1.16357421875,
144
- "step": 18
145
- },
146
- {
147
- "epoch": 0.42857142857142855,
148
- "eval_loss": 1.044921875,
149
- "eval_runtime": 2.1834,
150
- "eval_samples_per_second": 2.748,
151
- "eval_steps_per_second": 0.458,
152
- "step": 18
153
- },
154
- {
155
- "epoch": 0.4523809523809524,
156
- "grad_norm": 0.0762699693441391,
157
- "learning_rate": 7.525916851679529e-05,
158
- "loss": 1.06640625,
159
- "step": 19
160
- },
161
- {
162
- "epoch": 0.47619047619047616,
163
- "grad_norm": 0.05142414569854736,
164
- "learning_rate": 7.177854150011389e-05,
165
- "loss": 0.83837890625,
166
- "step": 20
167
- },
168
- {
169
- "epoch": 0.5,
170
- "grad_norm": 0.0549125075340271,
171
- "learning_rate": 6.817702470744477e-05,
172
- "loss": 0.98583984375,
173
- "step": 21
174
- },
175
- {
176
- "epoch": 0.5238095238095238,
177
- "grad_norm": 0.04964558407664299,
178
- "learning_rate": 6.448056714980767e-05,
179
- "loss": 1.0419921875,
180
- "step": 22
181
- },
182
- {
183
- "epoch": 0.5476190476190477,
184
- "grad_norm": 0.04620349779725075,
185
- "learning_rate": 6.071580188860955e-05,
186
- "loss": 0.9892578125,
187
- "step": 23
188
- },
189
- {
190
- "epoch": 0.5714285714285714,
191
- "grad_norm": 0.053955476731061935,
192
- "learning_rate": 5.690985414382668e-05,
193
- "loss": 1.12255859375,
194
- "step": 24
195
- },
196
- {
197
- "epoch": 0.5952380952380952,
198
- "grad_norm": 0.04355732351541519,
199
- "learning_rate": 5.3090145856173346e-05,
200
- "loss": 1.06640625,
201
- "step": 25
202
- },
203
- {
204
- "epoch": 0.6190476190476191,
205
- "grad_norm": 0.04457163065671921,
206
- "learning_rate": 4.9284198111390456e-05,
207
- "loss": 1.1181640625,
208
- "step": 26
209
- },
210
- {
211
- "epoch": 0.6190476190476191,
212
- "eval_loss": 1.0146484375,
213
- "eval_runtime": 2.2105,
214
- "eval_samples_per_second": 2.714,
215
- "eval_steps_per_second": 0.452,
216
- "step": 26
217
- },
218
- {
219
- "epoch": 0.6428571428571429,
220
- "grad_norm": 0.048139654099941254,
221
- "learning_rate": 4.551943285019234e-05,
222
- "loss": 1.0576171875,
223
- "step": 27
224
- },
225
- {
226
- "epoch": 0.6666666666666666,
227
- "grad_norm": 0.045640308409929276,
228
- "learning_rate": 4.182297529255525e-05,
229
- "loss": 1.025390625,
230
- "step": 28
231
- },
232
- {
233
- "epoch": 0.6904761904761905,
234
- "grad_norm": 0.04443604126572609,
235
- "learning_rate": 3.822145849988612e-05,
236
- "loss": 1.1689453125,
237
- "step": 29
238
- },
239
- {
240
- "epoch": 0.7142857142857143,
241
- "grad_norm": 0.040993694216012955,
242
- "learning_rate": 3.474083148320469e-05,
243
- "loss": 1.07568359375,
244
- "step": 30
245
- },
246
- {
247
- "epoch": 0.7380952380952381,
248
- "grad_norm": 0.03656896576285362,
249
- "learning_rate": 3.1406172239924584e-05,
250
- "loss": 0.955322265625,
251
- "step": 31
252
- },
253
- {
254
- "epoch": 0.7619047619047619,
255
- "grad_norm": 0.04126700386404991,
256
- "learning_rate": 2.8241507066307104e-05,
257
- "loss": 1.0849609375,
258
- "step": 32
259
- },
260
- {
261
- "epoch": 0.7857142857142857,
262
- "grad_norm": 0.043409474194049835,
263
- "learning_rate": 2.5269637447446348e-05,
264
- "loss": 1.04052734375,
265
- "step": 33
266
- },
267
- {
268
- "epoch": 0.8095238095238095,
269
- "grad_norm": 0.03532347083091736,
270
- "learning_rate": 2.2511975772046403e-05,
271
- "loss": 0.845703125,
272
- "step": 34
273
- },
274
- {
275
- "epoch": 0.8095238095238095,
276
- "eval_loss": 1.00341796875,
277
- "eval_runtime": 2.2033,
278
- "eval_samples_per_second": 2.723,
279
- "eval_steps_per_second": 0.454,
280
- "step": 34
281
- },
282
- {
283
- "epoch": 0.8333333333333334,
284
- "grad_norm": 0.04271344095468521,
285
- "learning_rate": 1.9988391055670233e-05,
286
- "loss": 1.0478515625,
287
- "step": 35
288
- },
289
- {
290
- "epoch": 0.8571428571428571,
291
- "grad_norm": 0.044091783463954926,
292
- "learning_rate": 1.771706578402711e-05,
293
- "loss": 1.02978515625,
294
- "step": 36
295
- },
296
- {
297
- "epoch": 0.8809523809523809,
298
- "grad_norm": 0.03613298386335373,
299
- "learning_rate": 1.5714364907746536e-05,
300
- "loss": 0.956756591796875,
301
- "step": 37
302
- },
303
- {
304
- "epoch": 0.9047619047619048,
305
- "grad_norm": 0.03579853102564812,
306
- "learning_rate": 1.3994717932533891e-05,
307
- "loss": 0.92950439453125,
308
- "step": 38
309
- },
310
- {
311
- "epoch": 0.9285714285714286,
312
- "grad_norm": 0.03729414567351341,
313
- "learning_rate": 1.257051495425121e-05,
314
- "loss": 1.02392578125,
315
- "step": 39
316
- },
317
- {
318
- "epoch": 0.9523809523809523,
319
- "grad_norm": 0.03735141083598137,
320
- "learning_rate": 1.1452017387992552e-05,
321
- "loss": 0.94476318359375,
322
- "step": 40
323
- },
324
- {
325
- "epoch": 0.9761904761904762,
326
- "grad_norm": 0.038969486951828,
327
- "learning_rate": 1.064728403435312e-05,
328
- "loss": 1.0830078125,
329
- "step": 41
330
- },
331
- {
332
- "epoch": 1.0,
333
- "grad_norm": 0.04045809060335159,
334
- "learning_rate": 1.0162113015586309e-05,
335
- "loss": 1.03173828125,
336
- "step": 42
337
- },
338
- {
339
- "epoch": 1.0,
340
- "eval_loss": 0.99658203125,
341
- "eval_runtime": 2.2148,
342
- "eval_samples_per_second": 2.709,
343
- "eval_steps_per_second": 0.452,
344
- "step": 42
345
- }
346
- ],
347
- "logging_steps": 1.0,
348
- "max_steps": 42,
349
- "num_input_tokens_seen": 0,
350
- "num_train_epochs": 1,
351
- "save_steps": 0,
352
- "stateful_callbacks": {
353
- "TrainerControl": {
354
- "args": {
355
- "should_epoch_stop": false,
356
- "should_evaluate": false,
357
- "should_log": false,
358
- "should_save": true,
359
- "should_training_stop": true
360
- },
361
- "attributes": {}
362
- }
363
- },
364
- "total_flos": 3.07543833476019e+18,
365
- "train_batch_size": 1,
366
- "trial_name": null,
367
- "trial_params": null
368
- }