Text Generation
Transformers
Safetensors
Portuguese
qwen3
text-generation-inference
conversational
Eval Results (legacy)
nicholasKluge commited on
Commit
283a303
·
verified ·
1 Parent(s): df6cc65

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +575 -0
README.md ADDED
@@ -0,0 +1,575 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - pt
4
+ license: apache-2.0
5
+ library_name: transformers
6
+ tags:
7
+ - text-generation-inference
8
+ datasets:
9
+ - Polygl0t/gigaverbo-v2-sft
10
+ - Polygl0t/gigaverbo-v2-preferences
11
+ metrics:
12
+ - perplexity
13
+ pipeline_tag: text-generation
14
+ widget:
15
+ - text: "<|im_start|>user\nQual é a capital de Portugal?<|im_end|><|im_start|>assistant\n"
16
+ example_title: Exemplo
17
+ - text: "<|im_start|>user\nEscreva um poema sobre a floresta amazônica.<|im_end|><|im_start|>assistant\n"
18
+ example_title: Exemplo
19
+ - text: "<|im_start|>user\nListe três benefícios da energia solar.<|im_end|><|im_start|>assistant\n"
20
+ example_title: Exemplo
21
+ inference:
22
+ parameters:
23
+ repetition_penalty: 1.2
24
+ temperature: 0.1
25
+ top_k: 50
26
+ top_p: 1.0
27
+ max_new_tokens: 150
28
+ co2_eq_emissions:
29
+ emissions: 2550
30
+ source: CodeCarbon
31
+ training_type: post-training
32
+ geographical_location: Germany
33
+ hardware_used: NVIDIA A100-SXM4-80GB
34
+ model-index:
35
+ - name: Tucano2-qwen-1.5B-Think
36
+ results:
37
+ - task:
38
+ type: text-generation
39
+ name: Text Generation
40
+ dataset:
41
+ name: ARC Challenge (Portuguese)
42
+ type: Polygl0t/ARC-poly
43
+ split: test
44
+ args:
45
+ num_few_shot: 5
46
+ metrics:
47
+ - type: acc_norm
48
+ value: 42.82
49
+ name: accuracy (normalized)
50
+ source:
51
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
52
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
53
+ - task:
54
+ type: text-generation
55
+ name: Text Generation
56
+ dataset:
57
+ name: HellaSwag (Portuguese)
58
+ type: Polygl0t/HellaSwag-poly
59
+ split: validation
60
+ args:
61
+ num_few_shot: 5
62
+ metrics:
63
+ - type: acc_norm
64
+ value: 54.95
65
+ name: accuracy (normalized)
66
+ source:
67
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
68
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
69
+ - task:
70
+ type: text-generation
71
+ name: Text Generation
72
+ dataset:
73
+ name: Calame
74
+ type: Polygl0t/CALAME-PT
75
+ split: test
76
+ args:
77
+ num_few_shot: 5
78
+ metrics:
79
+ - type: acc
80
+ value: 11.13
81
+ name: accuracy
82
+ source:
83
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
84
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
85
+ - task:
86
+ type: text-generation
87
+ name: Text Generation
88
+ dataset:
89
+ name: Lambada (Portuguese)
90
+ type: Polygl0t/LAMBADA-poly
91
+ split: test
92
+ args:
93
+ num_few_shot: 5
94
+ metrics:
95
+ - type: acc
96
+ value: 26.7
97
+ name: accuracy
98
+ source:
99
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
100
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
101
+ - task:
102
+ type: text-generation
103
+ name: Text Generation
104
+ dataset:
105
+ name: Global PIQA (por_latn_braz)
106
+ type: mrlbenchmarks/global-piqa-nonparallel
107
+ split: test
108
+ args:
109
+ num_few_shot: 5
110
+ metrics:
111
+ - type: acc_norm
112
+ value: 74
113
+ name: accuracy (normalized)
114
+ source:
115
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
116
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
117
+ - task:
118
+ type: text-generation
119
+ name: Text Generation
120
+ dataset:
121
+ name: MMLU (Portuguese)
122
+ type: Polygl0t/MMLU-poly
123
+ split: test
124
+ args:
125
+ num_few_shot: 5
126
+ metrics:
127
+ - type: acc
128
+ value: 43.3
129
+ name: accuracy
130
+ source:
131
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
132
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
133
+ - task:
134
+ type: text-generation
135
+ name: Text Generation
136
+ dataset:
137
+ name: BELEBELE (Portuguese)
138
+ type: facebook/belebele
139
+ split: test
140
+ args:
141
+ num_few_shot: 5
142
+ metrics:
143
+ - type: acc_norm
144
+ value: 67.67
145
+ name: accuracy (normalized)
146
+ source:
147
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
148
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
149
+ - task:
150
+ type: text-generation
151
+ name: Text Generation
152
+ dataset:
153
+ name: BLUEX (No Images)
154
+ type: eduagarcia-temp/BLUEX_without_images
155
+ split: train
156
+ args:
157
+ num_few_shot: 3
158
+ metrics:
159
+ - type: acc
160
+ value: 39.22
161
+ name: accuracy
162
+ source:
163
+ url: https://github.com/eduagarcia/lm-evaluation-harness-pt
164
+ name: Open Portuguese LLM Leaderboard
165
+ - task:
166
+ type: text-generation
167
+ name: Text Generation
168
+ dataset:
169
+ name: ENEM Challenge (No Images)
170
+ type: eduagarcia/enem_challenge
171
+ split: train
172
+ args:
173
+ num_few_shot: 3
174
+ metrics:
175
+ - type: acc
176
+ value: 39.89
177
+ name: accuracy
178
+ source:
179
+ url: https://github.com/eduagarcia/lm-evaluation-harness-pt
180
+ name: Open Portuguese LLM Leaderboard
181
+ - task:
182
+ type: text-generation
183
+ name: Text Generation
184
+ dataset:
185
+ name: OAB Exams
186
+ type: eduagarcia/oab_exams
187
+ split: train
188
+ args:
189
+ num_few_shot: 3
190
+ metrics:
191
+ - type: acc
192
+ value: 34.26
193
+ name: accuracy
194
+ source:
195
+ url: https://github.com/eduagarcia/lm-evaluation-harness-pt
196
+ name: Open Portuguese LLM Leaderboard
197
+ - task:
198
+ type: text-generation
199
+ name: Text Generation
200
+ dataset:
201
+ name: IFEval (Portuguese)
202
+ type: Polygl0t/IFEval-PT
203
+ split: train
204
+ args:
205
+ num_few_shot: 0
206
+ metrics:
207
+ - type: ifeval_pt_inst_level_loose_acc
208
+ value: 45.12
209
+ name: accuracy (loose)
210
+ - type: ifeval_pt_prompt_level_loose_acc
211
+ value: 36.67
212
+ name: accuracy (loose)
213
+ source:
214
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
215
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
216
+ - task:
217
+ type: text-generation
218
+ name: Text Generation
219
+ dataset:
220
+ name: GSM8K (Portuguese)
221
+ type: Polygl0t/gsm8k-pt
222
+ split: test
223
+ args:
224
+ num_few_shot: 0
225
+ metrics:
226
+ - type: flexible-extract
227
+ value: 22.83
228
+ name: accuracy (flexible-extract)
229
+ source:
230
+ url: https://github.com/Nkluge-correa/lm-evaluation-harness
231
+ name: Language Model Evaluation Harness (branch=polyglot_harness_portuguese)
232
+ base_model: Polygl0t/Tucano2-qwen-1.5B-Base
233
+ ---
234
+
235
+ # Tucano2-qwen-1.5B-Think
236
+
237
+ <img src="./logo.png" alt="An illustration of a Tucano bird showing vibrant colors like yellow, orange, blue, green, and black." height="200">
238
+
239
+ ## Model Summary
240
+
241
+ **[Tucano2-qwen-1.5B-Think](https://huggingface.co/Polygl0t/Tucano2-qwen-1.5B-Think)** is an instruction-tuned Portuguese language model built on top of **Tucano2-qwen-0.5B-Base**. It has been trained using a combination of one round of supervised fine-tuning (SFT) and one round of Anchored Preference Optimization (APO).
242
+
243
+ Tucano2-qwen-1.5B-Think is a reasoning model, which means it has been fine-tuned to generate CoT-style (Chain-of-Thought) traces in its responses. These reasoning traces are always encapsulated within the special tokens `<think>` and `</think>`.
244
+
245
+ **All datasets, source code, and training recipes used to develop the Tucano2 series are fully open and reproducible.**
246
+
247
+ ## Details
248
+
249
+ - **Architecture:** a Transformer-based model pre-trained via causal language modeling
250
+ - **Size:** 490,799,104 parameters
251
+ - **Context length:** 4,096 tokens
252
+ - **Dataset(s):**
253
+ - [Polygl0t/gigaverbo-v2-sft](https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft)
254
+ - [Polygl0t/gigaverbo-v2-preferences](https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-preferences)
255
+ - **Training time**: ~ 1.5 hours
256
+ - **Emissions:** 1.23 KgCO2 (Germany)
257
+ - **Total energy consumption:** 2.66 kWh
258
+
259
+ This repository has the [source code](https://github.com/Polygl0t/Polygl0t) used to train this model. The full configuration used for training is available in the following config files:
260
+
261
+ - Single stage Supervised Fine-Tuning (linear warmup with cosine decay): [training_config_sft.yaml](training_config_sft.yaml)
262
+ - Single stage Anchored Preference Optimization (linear warmup with cosine decay): [training_config_apo.yaml](training_config_apo.yaml)
263
+ - Training Logs (loss, lr, rewards, etc.): [train_logs_apo.parquet](train_logs_apo.parquet), [train_logs_sft.parquet](train_logs_sft.parquet)
264
+
265
+ <details>
266
+ <summary><b>SFT Loss Curve</b></summary>
267
+
268
+ ![SFT Loss Curve](./.plots/sft_loss.png)
269
+
270
+ </details>
271
+
272
+ <details>
273
+ <summary><b>APO Rewards</b></summary>
274
+
275
+ ![APO Rewards](./.plots/apo_reward.png)
276
+
277
+ </details>
278
+
279
+ ## Intended Uses
280
+
281
+ The primary intended use Tucano2-qwen-1.5B-Think is to serve as foundations for research and development involving Portuguese language modeling. You may also fine-tune and adapt Tucano2-qwen-1.5B-Think for deployment if your use follows the Apache 2.0 license. If you decide to use Tucano2-qwen-1.5B-Think as a basis for your fine-tuned model, please conduct your own risk and bias assessment.
282
+
283
+ ## Basic usage
284
+
285
+ ```python
286
+ from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
287
+ import torch
288
+
289
+ # Load model and tokenizer
290
+ model_id = "Polygl0t/Tucano2-qwen-1.5B-Think"
291
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
292
+ model = AutoModelForCausalLM.from_pretrained(
293
+ model_id,
294
+ device_map="auto"
295
+ )
296
+
297
+ # Configure generation parameters
298
+ generation_config = GenerationConfig(
299
+ do_sample=True,
300
+ temperature=0.1,
301
+ top_k=50,
302
+ top_p=1.0,
303
+ repetition_penalty=1.2,
304
+ max_new_tokens=150,
305
+ pad_token_id=tokenizer.eos_token_id,
306
+ )
307
+
308
+ # Prepare chat messages
309
+ messages = [
310
+ {"role": "user", "content": "Qual é a capital de Angola?"}
311
+ ]
312
+
313
+ # Apply chat template and generate
314
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
315
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
316
+
317
+ with torch.no_grad():
318
+ outputs = model.generate(**inputs, generation_config=generation_config)
319
+
320
+ # Decode and print response
321
+ response = tokenizer.decode(outputs[0][len(inputs.input_ids[0]):], skip_special_tokens=True)
322
+ print(f"🤖 {response}")
323
+ ```
324
+
325
+ ## Limitations
326
+
327
+ Like almost all other language models trained on large text datasets scraped from the web, the Tucano2-qwen-1.5B-Think shows behavior that does not make it an out-of-the-box solution to many real-world applications, especially those requiring factual, reliable, and nontoxic text generation. Tucano2-qwen-1.5B-Think is subject to the following:
328
+
329
+ - **Hallucinations:** Tucano2-qwen-1.5B-Think can produce content that can be mistaken as true facts, but are misleading or entirely false, i.e., hallucination.
330
+
331
+ - **Biases and Toxicity:** Tucano2-qwen-1.5B-Think inherits the social and historical stereotypes from the data used to train it. Given these biases, the model can produce toxic content, i.e., harmful, offensive, or detrimental to individuals, groups, or communities.
332
+
333
+ - **Language Limitations:** Tucano2-qwen-1.5B-Think is primarily designed to interact with Portuguese. Other languages might challenge its comprehension, leading to potential misinterpretations or errors in response.
334
+
335
+ - **Repetition and Verbosity:** Tucano2-qwen-1.5B-Think may get stuck on repetition loops (especially if the repetition penalty during generations is set to a meager value) or produce verbose responses unrelated to the prompt it was given.
336
+
337
+ Hence, even though Tucano2-qwen-1.5B-Think is released with a permissive license, we urge users to perform their risk analysis on them if they intend to use them for real-world applications.
338
+
339
+ ## Evaluations
340
+
341
+ We evaluated Tucano2-qwen-1.5B-Think on a variety of benchmarks covering different tasks and difficulty levels. The results are summarized below.
342
+
343
+ ### Inference Samples
344
+
345
+ Here we provide some inference samples generated by Tucano2-qwen-1.5B-Think.
346
+
347
+ <details>
348
+ <summary><b>Inference Samples on Various Tasks</b></summary>
349
+
350
+ ### Cooking
351
+
352
+ **Prompt:**
353
+
354
+ <|im_start|>user
355
+ Você pode me dar uma receita simples de Bolo de Brigadeiro?<|im_end|>
356
+ <|im_start|>assistant
357
+
358
+ **Response:**
359
+
360
+ <think>
361
+ O usuário está me pedindo uma receita simplificada de bolo de brigadeiro. Primeiro, vou analisar a solicitação para entender exatamente o que ele precisa. Preciso garantir que a receita seja fácil e rápida de seguir, então decido omitir alguns passos técnicos mais complexos. Agora preciso considerar os ingredientes básicos necessários para fazer um bom bolo de brigadeiro. Vou listar todos eles na minha mente: farinha de trigo, açúcar refinado, ovos grandes, óleo vegetal, leite condensado, chocolate amargo picado e gelatina em pó sem sabor. Devo organizar esses itens numa lista clara e concisa. Além disso, é importante mencionar as etapas principais da preparação do recheio (gelatina dissolvida no leite) e dos dois bolos individuais (farinha peneirada misturada com água morna). Também devo sugerir como untar e enfarinhar formas retangulares pequenas antes de colocar os bolos dentro delas. Finalmente, adicionarei instruções rápidas sobre assar nos forno pré-aquecido e resfriá-los completamente após retirar das formas. Antes de concluir, ofereço ao usuário a opção de pedir dicas extras caso precise. Isso mostra disponibilidade e ajuda quem possa estar buscando informações adicionais.
362
+ </think>
363
+
364
+ Claro! Aqui vai uma versão bem simplificada de **Bolo de Brigadeiro** que você pode preparar rapidamente:
365
+
366
+ ### Ingredientes:
367
+ - 2 xícaras de farinha de trigo
368
+ - 1 xícara de açúcar refinado
369
+ - 4 ovos grandes
370
+ - ½ colher de chá de fermento químico incolor
371
+ - ⅓ copo de óleo vegetal
372
+ - 3 copos de leite condensado
373
+ - ₁ lata de chocolate amargo picado (aproximadamente 240g)
374
+ - Gelatina em pó sem sabor (1 tablete)
375
+
376
+ ### Modo de Preparo:
377
+ 1. Prepare a forma: Untie e dusted duas formas retangulares pequenos ou use papel manteiga.
378
+ 2. Pré-aqueça o forno a 180°C.
379
+ 3. Em uma panela pequena, derreta metade da gelatina em pó num pouco de água quente. Reserve.
380
+ 4. Misture a outra metade da gelatina não derretida junto com todo o resto da geléia nas formas untadas. Deixe descansar por cerca de 5 minutos até ficar consistente.
381
+ 5. Na batedeira, batam os ovos com o açúcar até obter uma mistura cremosa. Adicione gradualmente a farinha peneirada e mexendo apenas suficiente para combinar.
382
+ 6. Acrescente o fermento químico aos poucos à massa já formada.
383
+ 7. Gradualmente, incorpore o óleo vegetal às massas secas.
384
+ 8. Por último, junte o leite condensado e o chocolate amargo picado. Mexer suavemente só pra misturar tudo uniformemente.
385
+ 9. Preencha cada cavidade das formas preparadas anteriormente com esta mistura.
386
+ 10. Asse durante 25-30 minutos aproximadamente, verificando se estão dourados pelo lado externo usando palitos finos.
387
+ 11. Retire do forno e deixe esfriando totalmente dentro das formas.
388
+ 12. Transfira cuidadosamente para bandejas de gelo cortáveis enquanto ainda estiverem quentes mas manejáveis.
389
+ 13. Sirva fresco logo depois de desconectar do congelador.
390
+
391
+ E pronto! Um delicioso Bolinho de Brigadeiros caseiro rapidinho. Espero que goste! Qualquer coisa, estou aqui pra ajudar!<|im_end|>
392
+
393
+ ---
394
+
395
+ ### Math
396
+
397
+ **Prompt:**
398
+
399
+ <|im_start|>user
400
+ Como eu posso resolver o seguinte problema: 2x + 3 = 11?<|im_end|>
401
+ <|im_start|>assistant
402
+
403
+ **Response:**
404
+
405
+ <think>
406
+ O usuário está me perguntando como resolve-se a equação 2x + 3 = 11. Primeiro, vou analisar cuidadosamente cada elemento da equação. Preciso isolar x no lado esquerdo para descobrir seu valor. Vou começar subtraindo 3 de ambos os lados para eliminar esse termo constante do lado esquerdo. Isso dá duas vezes X igual ao número menos três e onze juntos. Agora tenho que dividir tudo por dois porque quero deixar só X multiplicado pelo coeficiente 2. Deixe-me fazer essa divisão agora. Dividindo 20/2 obtenho 10, então minha solução é x=10. Mas antes de confirmar isso, preciso revisitar todos os passos pra garantir nada foi deixado escapar ou cometido erro algum na manipulação algébrica. Sim, parece correto! Então posso dizer com confiança que a resposta correta é realmente 10.
407
+ </think>
408
+
409
+ Claro! Vamos resolver a equação \( 2x + 3 = 11 \) passo a passo.
410
+
411
+ ### Passo 1: Isolar o termo com a variável \( x \)
412
+ Primeiro, vamos remover o termo independente (que não contém \( x \)) do lado onde estamos guardando nossa incógnita. Para isso, subtraímos 3 de ambos os lados da equação:
413
+
414
+ \[
415
+ 2x + 3 - 3 = 11 - 3 \\
416
+ 2x = 8
417
+ \]
418
+
419
+ Agora temos apenas \( 2x = 8 \).
420
+
421
+ ### Passo 2: Resolver para \( x \)
422
+ Para encontrar o valor de \( x \), precisamos isolá-lo dividindo ambos os lados pela constante que multiplica \( x \):
423
+
424
+ \[
425
+ \frac{2x}{2} = \frac{8}{2}\quad (\text{dividimos ambos os lados por } 2)\\
426
+ x=4
427
+ \]
428
+
429
+ Portanto, a solução para a equação \( 2x + 3 = 11 \) é \( x = 4 \).
430
+
431
+ Se você tiver alguma outra dúvida ou quiser ver outro exemplo, estou aqui para ajudar!<|im_end|>
432
+
433
+ ---
434
+
435
+ ### Rewriting
436
+
437
+ **Prompt:**
438
+
439
+ <|im_start|>user
440
+ Remova todas as informações irrelevantes da frase a seguir.
441
+
442
+ "O novo restaurante que abriu no centro da cidade, de propriedade do primo de Jorge, que costumava ser chef de um restaurante em Ipanema, serve uma variedade de cozinhas de todo o mundo."<|im_end|>
443
+ <|im_start|>assistant
444
+
445
+ **Response:**
446
+
447
+ <think>
448
+ O usuário está me pedindo para remover toda informação desnecessária da frase fornecida. Primeiro, vou ler cuidadosamente a frase completa: 'O novo restaurante que abriu no centro da cidade, de propriedade do primo de Jorge, que costumava ser chefe de um restaurante em Ipanema, serve uma variedade de cozinhas de todo o mundo.' Agora preciso identificar quais partes são realmente importantes e relevantes ao objetivo principal dessa pergunta — entender qual tipo específico de cozinha esse restaurante oferece. Vou começar eliminando detalhes secundários ou redundâncias.
449
+ </think>
450
+
451
+ Entendi! Vamos focar nas informações essenciais sobre o tipo de culinária oferecida pelo restaurante:
452
+
453
+ > O novo restaurante no centro da cidade serve uma variedade de cozinhas de todo o mundo.
454
+
455
+ Se você quiser alguma outra coisa específica, estou aqui pra ajudar!<|im_end|>
456
+
457
+ ---
458
+
459
+ ### History
460
+
461
+ **Prompt:**
462
+
463
+ <|im_start|>user
464
+ Liste os principais eventos envolvendo a Revolução Farroupilha.<|im_end|>
465
+ <|im_start|>assistant
466
+
467
+ **Response:**
468
+
469
+ <think>
470
+ O usuário está me perguntando para listar os principais eventos envolvidos na Revolução Farroupilha. Primeiro, vou analisar exatamente o que ele quer saber: são fatos cronológicos importantes e significativos dessa importante revolta gaúcha? Agora preciso lembrar dos detalhes históricos relevantes da época. Vou começar pelo início do movimento, pois isso dá contexto ao leitor sobre quando tudo começou. Devo mencionar as razões iniciais pelas quais surgiu esse conflito entre republicanos farrapos e autoridades imperiais. Depois disso, posso organizar minha lista em ordem temporal, destacando cada evento crucial. Preciso garantir que aborde tanto aspectos políticos quanto sociais-econômicos, já que esses foram elementos cruciais nos motivadores das revoltas locais. Também devo incluir datas aproximadas onde possível para facilitar a compreensão do tempo histórico. Finalmente, concluirei com uma nota geral sobre o impacto duradouro desta revolução no Rio Grande do Sul moderno. Isso ajudará o usuário a ter uma visão completa e contextualizada.
471
+ </think>
472
+
473
+ Claro! A Revolução Farroupilha, também conhecida como Guerra dos Farrapos ou Revolta dos Sargentos, foi um levante militar contra o Império Brasileiro ocorrido nas regiões sulistas atuais de Santa Catarina, Rio Grande do Sul e parte de São Paulo, durante o período imperial (1835-1845). Aqui estão alguns dos principais eventos:
474
+
475
+ 1. **Início da Rebelião**:
476
+ - Em março de 1831, eclodiu a primeira insurreição liderada por Bento Gonçalves, ex-soldado desertor do exército regular brasileiro. Inicialmente, esta rebelião visava derrubar o governo centralizado sob dom Pedro I. No entanto, após algumas vitórias militares, ela se transformou numa luta pela autonomia provincial.
477
+
478
+ 2. **Formação das Províncias Livres**:
479
+ - Com apoio popular crescente, grupos regionais formaram suas próprias forças armadas independentes chamadas "farrapos". Estas províncias livres incluíram Rio Grande (atual RS), Santa Catarina e Paraná. O líder máximo desses movimentos era Manuel Marques Lisboa, futuro marquês de Alegrete.
480
+
481
+ 3. **Conflito Militar Principal**:
482
+ - As hostilidades começaram oficialmente em junho de 1836, marcando o começo formal da guerra civil rio-grandense. Os farrapos enfrentaram tropas imperiais lideradas inicialmente pelos generais David Canabarro e Félix Zavala. Apesar das dificuldades logísticas e numéricas, eles conseguiram manter controle significativo sobre áreas estratégicas.
483
+
484
+ 4. **Negociação e Trégua Temporária**:
485
+ - Nos anos seguintes, houve períodos alternados de combates intensos seguidos por acordos negociados visando aliviar tensões temporariamente. Um exemplo notório é a assinatura do Tratado de Ponche Verde em outubro de 1841, embora tenha sido rompido rapidamente depois dele.
486
+
487
+ 5. **Finalização do Conflito**:
488
+ - Após várias campanhas bem-sucedidas conduzidas principalmente pelo general José Antônio Correia de Melo, finalmente resultou na queda definitiva de Porto Alegre às mãos imperiais em maio de 1844, encerrando assim efetivamente a maior rebeldia regionalista até então registrada dentro fronteiras brasileiras formais.
489
+
490
+ 6. **Impacto Político e Social**:
491
+ - Este conflito teve consequências profundas não só política mas socialmente. Resultou na criação legal da Província do Rio Grande do Sul através Lei Complementar nº 2, assinada logo após término bélico; além disto contribuiu decisivamente formação identidade nacional brasileira questionada muitos intelectuais contemporâneos revolucionários.
492
+
493
+ Esses pontos resumem brevemente alguns momentos chave daquela história rica e complexa. Espero que ajude você entender melhor este episódio fascinante da nossa herança cultural!<|im_end|>
494
+
495
+ ---
496
+
497
+ </details>
498
+
499
+ ## Evaluations
500
+
501
+ The table below compares the Tucano2 series against other base models of similar size. The NPM (Normalized Performance Metric) provides a balanced view of model performance across tasks, accounting for each task's inherent difficulty by normalizing each task's evaluation score by its respective random baseline.
502
+
503
+ <details>
504
+ <summary><b>Benchmarks Used</b></summary>
505
+
506
+ | **Benchmark** | **n-shot** | **Origin** | **Type** | **Baseline** | **Metric** |
507
+ | ----------------------- | ---------- | ---------- | --------------------- | ------------ | ------------------------ |
508
+ | **Easy Set** | | | | | |
509
+ | CALAME | 5-shot | Native | Next Token Prediction | 0 | `acc` |
510
+ | GlobalPIQA | 5-shot | Native | Commonsense Reasoning | 50 | `acc_norm` |
511
+ | LAMBADA | 5-shot | Translated | Next Token Prediction | 0 | `acc` |
512
+ | ARC-Challenge | 5-shot | Translated | Q&A | 25 | `acc_norm` |
513
+ | HellaSwag | 5-shot | Translated | Q&A | 25 | `acc_norm` |
514
+ | **Hard Set** | | | | | |
515
+ | ENEM | 3-shot | Native | Q&A | 20 | `acc` |
516
+ | BLUEX | 3-shot | Native | Q&A | 22.5 | `acc` |
517
+ | OAB Exams | 3-shot | Native | Q&A | 25 | `acc` |
518
+ | BELEBELE | 5-shot | Native | Q&A | 25 | `acc_norm` |
519
+ | MMLU | 5-shot | Translated | Q&A | 25 | `acc` |
520
+ | **Instruction Set** | | | | | |
521
+ | IFEval-PT (prompt) | 0-shot | Translated | Q&A | 0 | `prompt_level_loose_acc` |
522
+ | IFEval-PT (instruction) | 0-shot | Translated | Q&A | 0 | `inst_level_loose_acc` |
523
+ | GSM8K-PT | 0-shot | Translated | Math Problems | 0 | `flexible-extract` |
524
+
525
+ </details>
526
+
527
+ | | Aggregate NPM | NPM Easy | NPM Hard | NPM Instruction | BLUEX | ENEM | OAB | ARC Challenge | BELEBELE | CALAME | Global PIQA | HellaSwag | LAMBADA | MMLU | IFEval-PT (prompt) | IFEval-PT (instruction) | GSM8K-PT (flex) |
528
+ | -------------------------- | ------------- | -------- | -------- | --------------- | ------ | ------ | ------ | ------------- | -------- | ------ | ----------- | --------- | ------- | ------ | ------------------ | ----------------------- | --------------- |
529
+ | Qwen2.5-7B-Instruct | 58.9618 | 50.3498 | 59.372 | 72.6316 | 0.6579 | 0.739 | 0.5276 | 0.4957 | 0.88 | 0.4976 | 0.78 | 0.6419 | 0.6097 | 0.6447 | 0.7533 | 0.8023 | 0.6233 |
530
+ | Jurema-7B | 54.6559 | 58.1519 | 57.5188 | 44.0576 | 0.6342 | 0.7096 | 0.6497 | 0.5256 | 0.8844 | 0.6469 | 0.79 | 0.6732 | 0.7489 | 0.4991 | 0.47 | 0.5488 | 0.3029 |
531
+ | Gemma-3-Gaia-PT-BR-4b-it | 51.6711 | 49.6677 | 44.8169 | 66.4338 | 0.509 | 0.6452 | 0.4346 | 0.547 | 0.7889 | 0.5188 | 0.77 | 0.6221 | 0.5325 | 0.5149 | 0.7033 | 0.7767 | 0.5129 |
532
+ | SmolLM3-3B | 50.9191 | 51.1554 | 43.3707 | 63.1058 | 0.4868 | 0.6088 | 0.4223 | 0.5282 | 0.7844 | 0.604 | 0.71 | 0.6316 | 0.654 | 0.533 | 0.69 | 0.7488 | 0.4543 |
533
+ | Tucano2-qwen-3.7B-Think | 49.9686 | 52.8665 | 51.8393 | 42.0208 | 0.5341 | 0.6298 | 0.5376 | 0.547 | 0.8433 | 0.5467 | 0.81 | 0.6568 | 0.5381 | 0.611 | 0.3033 | 0.4116 | 0.5457 |
534
+ | Qwen2.5-3B-Instruct | 47.913 | 36.0891 | 51.4375 | 61.7453 | 0.5688 | 0.6865 | 0.4679 | 0.4171 | 0.84 | 0.4576 | 0.67 | 0.5844 | 0.3382 | 0.5822 | 0.6333 | 0.7 | 0.519 |
535
+ | Qwen2.5-1.5B-Instruct | 43.2133 | 41.4916 | 43.9841 | 44.798 | 0.5202 | 0.6179 | 0.4428 | 0.3974 | 0.76 | 0.5674 | 0.69 | 0.5021 | 0.5944 | 0.5191 | 0.42 | 0.5023 | 0.4216 |
536
+ | Tucano2-qwen-1.5B-Instruct | 42.6073 | 45.8889 | 44.779 | 33.5186 | 0.5285 | 0.627 | 0.4342 | 0.5026 | 0.7756 | 0.5438 | 0.71 | 0.5525 | 0.5905 | 0.5254 | 0.3433 | 0.4651 | 0.1971 |
537
+ | Llama-3.2-3B-Instruct | 39.8516 | 21.7714 | 44.243 | 62.6661 | 0.5202 | 0.5913 | 0.4497 | 0.4393 | 0.7856 | 0.0039 | 0.71 | 0.5583 | 0.0012 | 0.5214 | 0.6267 | 0.7023 | 0.551 |
538
+ | Tucano2-qwen-1.5B-Think | 30.0924 | 29.9038 | 28.0136 | 33.8713 | 0.3922 | 0.3989 | 0.3426 | 0.4282 | 0.6767 | 0.1113 | 0.74 | 0.5495 | 0.267 | 0.433 | 0.3367 | 0.4512 | 0.2283 |
539
+ | Qwen3-1.7B | 29.8345 | 9.679 | 34.8513 | 55.0655 | 0.541 | 0.606 | 0.4159 | 0.3667 | 0.6489 | 0.0014 | 0.56 | 0.4045 | 0.001 | 0.3056 | 0.65 | 0.7326 | 0.2694 |
540
+ | Tucano2-qwen-0.5B-Instruct | 29.7779 | 27.8097 | 31.5423 | 30.1179 | 0.4033 | 0.536 | 0.4073 | 0.3863 | 0.6233 | 0.3001 | 0.62 | 0.4783 | 0.3643 | 0.4146 | 0.3 | 0.4186 | 0.1849 |
541
+ | Qwen3-0.6B | 29.6678 | 19.0508 | 27.963 | 50.2042 | 0.4061 | 0.4591 | 0.3604 | 0.3239 | 0.6089 | 0.3931 | 0.49 | 0.372 | 0.3183 | 0.4112 | 0.55 | 0.6395 | 0.3166 |
542
+ | Qwen3-4B | 23.1764 | 16.489 | 2.3256 | 69.0736 | 0.0682 | 0.0602 | 0.0141 | 0.4308 | 0.8367 | 0.0048 | 0.65 | 0.4587 | 0.0004 | 0.2693 | 0.81 | 0.8558 | 0.4064 |
543
+ | Qwen2.5-0.5B-Instruct | 21.4918 | 20.9421 | 17.3804 | 29.2603 | 0.3018 | 0.3408 | 0.2934 | 0.2744 | 0.5067 | 0.4576 | 0.5 | 0.3774 | 0.3872 | 0.3954 | 0.31 | 0.4209 | 0.1469 |
544
+ | Tucano2-qwen-0.5B-Think | 19.2755 | 21.2555 | 12.5445 | 27.1936 | 0.3449 | 0.3198 | 0.2702 | 0.3274 | 0.3611 | 0.0949 | 0.68 | 0.4721 | 0.2086 | 0.3608 | 0.2767 | 0.393 | 0.1461 |
545
+ | Llama-3.2-1B-Instruct | 17.9105 | 6.8658 | 14.1275 | 42.6234 | 0.3004 | 0.3401 | 0.3084 | 0.3282 | 0.4156 | 0.0092 | 0.49 | 0.4362 | 0.0016 | 0.3515 | 0.4433 | 0.5698 | 0.2656 |
546
+ | Tucano-2b4-Instruct | 9.3564 | 14.456 | 1.5293 | 13.902 | 0.2587 | 0.2001 | 0.2674 | 0.3197 | 0.24 | 0.0048 | 0.67 | 0.4635 | 0.0004 | 0.2672 | 0.15 | 0.2465 | 0.0205 |
547
+ | Tucano-1b1-Instruct | 7.7272 | 12.0267 | 0.3183 | 12.9095 | 0.2295 | 0.1994 | 0.2533 | 0.3 | 0.2489 | 0 | 0.64 | 0.441 | 0 | 0.2559 | 0.1333 | 0.2372 | 0.0167 |
548
+ | TeenyTinyLlama-460m-Chat | 3.9618 | 6.58 | -3.25 | 11.6177 | 0.1725 | 0.1819 | 0.1973 | 0.2684 | 0.2289 | 0 | 0.59 | 0.3434 | 0 | 0.2697 | 0.1233 | 0.2047 | 0.0205 |
549
+
550
+ <details>
551
+ <summary><b>Performance Comparison</b></summary>
552
+
553
+ Below, we compare the performance of Tucano2-qwen-1.5B-Think with Qwen3-1.7B, a strong baseline in the 1.5B parameter range. The percentages represent the absolute difference in performance between the two models on each benchmark. All other plots can be found in the [.plots](https://huggingface.co/Polygl0t/Tucano2-qwen-1.5B-Think/tree/main/.plots/) folder.
554
+
555
+ **Tucano2-qwen-1.5B-Think vs Qwen3-1.7B**
556
+
557
+ ![Performance Comparison](./.plots/model_comparison.png)
558
+
559
+ </details>
560
+
561
+ ## Cite as 🤗
562
+
563
+ ```latex
564
+
565
+ ```
566
+
567
+ ## Aknowlegments
568
+
569
+ Polyglot is a project funded by the Federal Ministry of Education and Research (BMBF) and the Ministry of Culture and Science of the State of North Rhine-Westphalia (MWK) as part of TRA Sustainable Futures (University of Bonn) and the Excellence Strategy of the federal and state governments.
570
+
571
+ We also gratefully acknowledge the granted access to the [Marvin cluster](https://www.hpc.uni-bonn.de/en/systems/marvin) hosted by [University of Bonn](https://www.uni-bonn.de/en) along with the support provided by its High Performance Computing & Analytics Lab.
572
+
573
+ ## License
574
+
575
+ Tucano2-qwen-1.5B-Think is licensed under the Apache License, Version 2.0. For more details, see the [LICENSE](LICENSE) file.