SAnocha commited on
Commit
e803b65
·
verified ·
1 Parent(s): ba19d7a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +5 -0
README.md CHANGED
@@ -115,6 +115,11 @@ Note:
115
  - Tamil data from Sangraha is published [here](https://huggingface.co/datasets/ai4bharat/sangraha). The paper can be found [here](https://arxiv.org/abs/2403.06350).
116
  - Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)
117
 
 
 
 
 
 
118
  ## Call for Contributions
119
  We encourage researchers, developers, and language enthusiasts to actively contribute to the enhancement and expansion of SEA-LION. Contributions can involve identifying and reporting bugs, sharing pre-training, instruction, and preference data, improving documentation usability, proposing and implementing new model evaluation tasks and metrics, or training versions of the model in additional Southeast Asian languages. Join us in shaping the future of SEA-LION by sharing your expertise and insights to make these models more accessible, accurate, and versatile. Please check out our GitHub for further information on the call for contributions.
120
 
 
115
  - Tamil data from Sangraha is published [here](https://huggingface.co/datasets/ai4bharat/sangraha). The paper can be found [here](https://arxiv.org/abs/2403.06350).
116
  - Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)
117
 
118
+ ### Lineage & Versioning
119
+
120
+ The training counts and dataset mixture details reported in this model card reflect the exact constructed training pool consumed during this specific model run. Figures may differ slightly from public dataset releases, which represent downloadable open-source subsets of the broader corpus.
121
+
122
+
123
  ## Call for Contributions
124
  We encourage researchers, developers, and language enthusiasts to actively contribute to the enhancement and expansion of SEA-LION. Contributions can involve identifying and reporting bugs, sharing pre-training, instruction, and preference data, improving documentation usability, proposing and implementing new model evaluation tasks and metrics, or training versions of the model in additional Southeast Asian languages. Join us in shaping the future of SEA-LION by sharing your expertise and insights to make these models more accessible, accurate, and versatile. Please check out our GitHub for further information on the call for contributions.
125