|
Download README.md from ChrisToukmaji/focus_bur_mpt_focus_trained: direct link, hf CLI and curl.
- Browser
- Download file 2.34 kB
-
https://huggingface.co/ChrisToukmaji/focus_bur_mpt_focus_trained/resolve/main/README.md
- Command line
-
hf download hf://ChrisToukmaji/focus_bur_mpt_focus_trained/README.md
-
curl -L -o README.md https://huggingface.co/ChrisToukmaji/focus_bur_mpt_focus_trained/resolve/main/README.md
2.34 kB
| base_model: final_models/focus_bur_mpt_after_focus_reinit | |
| tags: | |
| - generated_from_trainer | |
| datasets: | |
| - mc4 | |
| model-index: | |
| - name: focus_bur_mpt_focus_trained | |
| results: [] | |
| <!-- This model card has been generated automatically according to the information the Trainer had access to. You | |
| should probably proofread and complete it, then remove this comment. --> | |
| # Paper and Citation | |
| Paper: [Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages | |
| ](https://arxiv.org/abs/2506.19187) | |
| ``` | |
| @misc{toukmaji2025prompttranslatefinetunereinitialize, | |
| title={Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages}, | |
| author={Christopher Toukmaji and Jeffrey Flanigan}, | |
| year={2025}, | |
| eprint={2506.19187}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={https://arxiv.org/abs/2506.19187}, | |
| } | |
| ``` | |
| # focus_bur_mpt_focus_trained | |
| This model is a fine-tuned version of [final_models/focus_bur_mpt_after_focus_reinit](https://huggingface.co/final_models/focus_bur_mpt_after_focus_reinit) on the mc4 my dataset. | |
| It achieves the following results on the evaluation set: | |
| - Loss: 2.0528 | |
| ## Model description | |
| More information needed | |
| ## Intended uses & limitations | |
| More information needed | |
| ## Training and evaluation data | |
| More information needed | |
| ## Training procedure | |
| ### Training hyperparameters | |
| The following hyperparameters were used during training: | |
| - learning_rate: 0.0003 | |
| - train_batch_size: 1 | |
| - eval_batch_size: 1 | |
| - seed: 42 | |
| - distributed_type: multi-GPU | |
| - optimizer: Adam with betas=(0.9,0.95) and epsilon=1e-05 | |
| - lr_scheduler_type: cosine | |
| - lr_scheduler_warmup_steps: 2000 | |
| - num_epochs: 6.0 | |
| ### Training results | |
| | Training Loss | Epoch | Step | Validation Loss | | |
| |:-------------:|:-----:|:------:|:---------------:| | |
| | 2.3438 | 1.0 | 24415 | 2.1199 | | |
| | 1.4219 | 2.0 | 48830 | 2.0162 | | |
| | 2.375 | 3.0 | 73245 | 1.9148 | | |
| | 1.0312 | 4.0 | 97660 | 1.8230 | | |
| | 1.1094 | 5.0 | 122075 | 1.8094 | | |
| | 0.5508 | 6.0 | 146490 | 2.0528 | | |
| ### Framework versions | |
| - Transformers 4.44.0 | |
| - Pytorch 2.5.1+cu124 | |
| - Datasets 3.2.0 | |
| - Tokenizers 0.19.1 | |