use-natural-pt / README.md
victormuryn's picture
Upload README.md with huggingface_hub
da41da9 verified
|
Raw
History Blame Contribute Delete
3.48 kB
---
base_model:
- sentence-transformers/paraphrase-multilingual-mpnet-base-v2
language:
- multilingual
- ar
- bg
- ca
- cs
- da
- de
- el
- en
- es
- et
- fa
- fi
- fr
- gl
- gu
- he
- hi
- hr
- hu
- hy
- id
- it
- ja
- ka
- ko
- ku
- lt
- lv
- mk
- mn
- mr
- ms
- my
- nb
- nl
- pl
- pt
- ro
- sk
- sl
- sq
- sr
- sv
- th
- tr
- uk
- ur
- vi
license: apache-2.0
library_name: sentence-transformers
tags:
- sentence-transformers
- feature-extraction
- sentence-similarity
- transformers
language_bcp47:
- fr-ca
- pt-br
- zh-cn
- zh-tw
pipeline_tag: sentence-similarity
datasets:
- victormuryn/wsd-training-dataset
---
# use-natural-pt
This model is a fine-tuned version of [paraphrase-multilingual-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2), trained on the Ukrainian text corpus [UberText 2.0](https://lang.org.ua/en/ubertext/) on the raw corpus with pool targets, with no data augmentation applied. It is part of the [Ukrainian Sentence Embeddings](https://huggingface.co/collections/victormuryn/ukrainian-sentence-embeddings-use) collection, which explores the effect of different training strategies on sentence embedding quality for Ukrainian.
## Model Description
The model was fine-tuned using a contrastive objective on UberText 2.0.
## Collection Overview
| Model | Description |
| ----------------------- | ----------------------------------------------------- |
| use-natural-no-pt | Raw UberText 2.0, no augmentation, no pool targets |
| **use-natural-pt** | Raw UberText 2.0, no augmentation, pool targets |
| use-generated-no-pt | Generated augmentation, no pool targets |
| use-generated-pt | Generated augmentation, pool targets |
| use-translation-no-pt | Back-translation augmentation, no pool targets |
| use-translation-pt | Back-translation augmentation, pool targets |
| use-mask-no-pt | Masking augmentation, no pool targets |
| use-mask-pt | Masking augmentation, pool targets |
| use-dropout-no-pt | Dropout augmentation, no pool targets |
| use-dropout-pt | Dropout augmentation, pool targets |
| use-token-shuffle-no-pt | Token-shuffling augmentation, no pool targets |
| use-token-shuffle-pt | Token-shuffling augmentation, pool targets |
| use-combined-no-pt | Combined augmentation strategies, no pool targets |
| use-combined-pt | Combined augmentation strategies, pool targets |
| use-stochastic-no-pt | Markov-based stochastic augmentation, no pool targets |
| use-stochastic-pt | Markov-based stochastic augmentation, pool targets |
## Usage
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("victormuryn/use-natural-pt")
sentences = [
"Проводжає сина мати захищати рідний край",
"Хоч би малесеньку хатину він мріяв мати над Дніпром",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
```
## Training Details
- **Base model:** paraphrase-multilingual-mpnet-base-v2
- **Training corpus:** UberText 2.0
- **Augmentation:** None
- **Pool targets:** Yes
## Citation
To be added
## License
Apache 2.0