Text Ranking
sentence-transformers
Safetensors
xlm-roberta
cross-encoder
reranker
Generated from Trainer
dataset_size:50
loss:BinaryCrossEntropyLoss
Eval Results (legacy)
text-embeddings-inference
Instructions to use OloriBern/musique-climb-50 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use OloriBern/musique-climb-50 with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("OloriBern/musique-climb-50") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Upload climb model for musique (trained on 50 queries)
Browse files- README.md +29 -30
- config.json +1 -1
- eval/CrossEncoderClassificationEvaluator_validation_results.csv +3 -0
- model.safetensors +1 -1
- tokenizer_config.json +7 -0
README.md
CHANGED
|
@@ -6,7 +6,6 @@ tags:
|
|
| 6 |
- generated_from_trainer
|
| 7 |
- dataset_size:50
|
| 8 |
- loss:BinaryCrossEntropyLoss
|
| 9 |
-
base_model: BAAI/bge-reranker-v2-m3
|
| 10 |
pipeline_tag: text-ranking
|
| 11 |
library_name: sentence-transformers
|
| 12 |
metrics:
|
|
@@ -18,7 +17,7 @@ metrics:
|
|
| 18 |
- recall
|
| 19 |
- average_precision
|
| 20 |
model-index:
|
| 21 |
-
- name: CrossEncoder
|
| 22 |
results:
|
| 23 |
- task:
|
| 24 |
type: cross-encoder-binary-classification
|
|
@@ -31,13 +30,13 @@ model-index:
|
|
| 31 |
value: 0.964769647696477
|
| 32 |
name: Accuracy
|
| 33 |
- type: accuracy_threshold
|
| 34 |
-
value: 0.
|
| 35 |
name: Accuracy Threshold
|
| 36 |
- type: f1
|
| 37 |
value: 0.9446808510638298
|
| 38 |
name: F1
|
| 39 |
- type: f1_threshold
|
| 40 |
-
value: 0.
|
| 41 |
name: F1 Threshold
|
| 42 |
- type: precision
|
| 43 |
value: 0.9568965517241379
|
|
@@ -46,19 +45,19 @@ model-index:
|
|
| 46 |
value: 0.9327731092436975
|
| 47 |
name: Recall
|
| 48 |
- type: average_precision
|
| 49 |
-
value: 0.
|
| 50 |
name: Average Precision
|
| 51 |
---
|
| 52 |
|
| 53 |
-
# CrossEncoder
|
| 54 |
|
| 55 |
-
This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model
|
| 56 |
|
| 57 |
## Model Details
|
| 58 |
|
| 59 |
### Model Description
|
| 60 |
- **Model Type:** Cross Encoder
|
| 61 |
-
- **Base model:** [
|
| 62 |
- **Maximum Sequence Length:** 1024 tokens
|
| 63 |
- **Number of Output Labels:** 1 label
|
| 64 |
<!-- - **Training Dataset:** Unknown -->
|
|
@@ -90,11 +89,11 @@ from sentence_transformers import CrossEncoder
|
|
| 90 |
model = CrossEncoder("cross_encoder_model_id")
|
| 91 |
# Get scores for pairs of texts
|
| 92 |
pairs = [
|
| 93 |
-
['
|
| 94 |
-
['
|
| 95 |
-
[
|
| 96 |
-
['
|
| 97 |
-
['
|
| 98 |
]
|
| 99 |
scores = model.predict(pairs)
|
| 100 |
print(scores.shape)
|
|
@@ -102,13 +101,13 @@ print(scores.shape)
|
|
| 102 |
|
| 103 |
# Or rank different texts based on similarity to a single text
|
| 104 |
ranks = model.rank(
|
| 105 |
-
'
|
| 106 |
[
|
| 107 |
-
'
|
| 108 |
-
'
|
| 109 |
-
'
|
| 110 |
-
|
| 111 |
-
|
| 112 |
]
|
| 113 |
)
|
| 114 |
# [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]
|
|
@@ -150,12 +149,12 @@ You can finetune this model on your own dataset.
|
|
| 150 |
| Metric | Value |
|
| 151 |
|:----------------------|:-----------|
|
| 152 |
| accuracy | 0.9648 |
|
| 153 |
-
| accuracy_threshold | 0.
|
| 154 |
| f1 | 0.9447 |
|
| 155 |
-
| f1_threshold | 0.
|
| 156 |
| precision | 0.9569 |
|
| 157 |
| recall | 0.9328 |
|
| 158 |
-
| **average_precision** | **0.
|
| 159 |
|
| 160 |
<!--
|
| 161 |
## Bias, Risks and Limitations
|
|
@@ -183,11 +182,11 @@ You can finetune this model on your own dataset.
|
|
| 183 |
| type | string | string | float |
|
| 184 |
| details | <ul><li>min: 39 characters</li><li>mean: 76.6 characters</li><li>max: 114 characters</li></ul> | <ul><li>min: 148 characters</li><li>mean: 511.34 characters</li><li>max: 1394 characters</li></ul> | <ul><li>min: 0.0</li><li>mean: 0.32</li><li>max: 1.0</li></ul> |
|
| 185 |
* Samples:
|
| 186 |
-
| sentence_0 | sentence_1
|
| 187 |
-
|:--------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------|
|
| 188 |
-
| <code>
|
| 189 |
-
| <code>
|
| 190 |
-
| <code>
|
| 191 |
* Loss: [<code>BinaryCrossEntropyLoss</code>](https://sbert.net/docs/package_reference/cross_encoder/losses.html#binarycrossentropyloss) with these parameters:
|
| 192 |
```json
|
| 193 |
{
|
|
@@ -325,9 +324,9 @@ You can finetune this model on your own dataset.
|
|
| 325 |
### Training Logs
|
| 326 |
| Epoch | Step | validation_average_precision |
|
| 327 |
|:-----:|:----:|:----------------------------:|
|
| 328 |
-
| 1.0 | 13 | 0.
|
| 329 |
-
| 2.0 | 26 | 0.
|
| 330 |
-
| 3.0 | 39 | 0.
|
| 331 |
|
| 332 |
|
| 333 |
### Framework Versions
|
|
|
|
| 6 |
- generated_from_trainer
|
| 7 |
- dataset_size:50
|
| 8 |
- loss:BinaryCrossEntropyLoss
|
|
|
|
| 9 |
pipeline_tag: text-ranking
|
| 10 |
library_name: sentence-transformers
|
| 11 |
metrics:
|
|
|
|
| 17 |
- recall
|
| 18 |
- average_precision
|
| 19 |
model-index:
|
| 20 |
+
- name: CrossEncoder
|
| 21 |
results:
|
| 22 |
- task:
|
| 23 |
type: cross-encoder-binary-classification
|
|
|
|
| 30 |
value: 0.964769647696477
|
| 31 |
name: Accuracy
|
| 32 |
- type: accuracy_threshold
|
| 33 |
+
value: 0.09269625693559647
|
| 34 |
name: Accuracy Threshold
|
| 35 |
- type: f1
|
| 36 |
value: 0.9446808510638298
|
| 37 |
name: F1
|
| 38 |
- type: f1_threshold
|
| 39 |
+
value: 0.049161043018102646
|
| 40 |
name: F1 Threshold
|
| 41 |
- type: precision
|
| 42 |
value: 0.9568965517241379
|
|
|
|
| 45 |
value: 0.9327731092436975
|
| 46 |
name: Recall
|
| 47 |
- type: average_precision
|
| 48 |
+
value: 0.9828499641745869
|
| 49 |
name: Average Precision
|
| 50 |
---
|
| 51 |
|
| 52 |
+
# CrossEncoder
|
| 53 |
|
| 54 |
+
This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model trained using the [sentence-transformers](https://www.SBERT.net) library. It computes scores for pairs of texts, which can be used for text reranking and semantic search.
|
| 55 |
|
| 56 |
## Model Details
|
| 57 |
|
| 58 |
### Model Description
|
| 59 |
- **Model Type:** Cross Encoder
|
| 60 |
+
<!-- - **Base model:** [Unknown](https://huggingface.co/unknown) -->
|
| 61 |
- **Maximum Sequence Length:** 1024 tokens
|
| 62 |
- **Number of Output Labels:** 1 label
|
| 63 |
<!-- - **Training Dataset:** Unknown -->
|
|
|
|
| 89 |
model = CrossEncoder("cross_encoder_model_id")
|
| 90 |
# Get scores for pairs of texts
|
| 91 |
pairs = [
|
| 92 |
+
['Khosrowabad, in the city where Pouya Tajik was born, is found in what county?', 'Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward'],
|
| 93 |
+
['Khosrowabad, in the city where Pouya Tajik was born, is found in what county?', 'Frédéric Chopin. Chopin also endowed popular dance forms with a greater range of melody and expression. Chopin\'s mazurkas, while originating in the traditional Polish dance (the mazurek), differed from the traditional variety in that they were written for the concert hall rather than the dance hall; "it was Chopin who put the mazurka on the European musical map." The series of seven polonaises published in his lifetime (another nine were published posthumously), beginning with the Op. 26 pair (published 1836), set a new standard for music in the form. His waltzes were also written specifically for the salon recital rather than the ballroom and are frequently at rather faster tempos than their dance-floor equivalents.'],
|
| 94 |
+
['Margraviate of the country of the Botanical Garden of the place Josef Victor Rohon was educated is an instance of?', 'Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward'],
|
| 95 |
+
['Khosrowabad, in the city where Pouya Tajik was born, is found in what county?', "Frédéric Chopin. Numerous recordings of Chopin's works are available. On the occasion of the composer's bicentenary, the critics of The New York Times recommended performances by the following contemporary pianists (among many others): Martha Argerich, Vladimir Ashkenazy, Emanuel Ax, Evgeny Kissin, Murray Perahia, Maurizio Pollini and Krystian Zimerman. The Warsaw Chopin Society organizes the Grand prix du disque de F. Chopin for notable Chopin recordings, held every five years."],
|
| 96 |
+
['Margraviate of the country of the Botanical Garden of the place Josef Victor Rohon was educated is an instance of?', 'Arignar Anna Zoological Park. The butterfly house, constructed at a cost of ₹6 million, has more than 25 host plants and landscaped habitats, such as bushes, lianas, streams, waterfall and rock - gardens, that attract many species of butterflies such as the common Mormon, crimson rose, mottled emigrant, blue tiger, evening brown and lime butterfly. A network of ponds interconnected by streams maintains humidity in the area. The park covers an area of 5 acres. The butterfly garden with an insect museum at the entrance is set up by the Tamil Nadu Agricultural University (TNAU), Coimbatore. The insect museum has been planned with an exhibit area comprising insect exhibits representing the most common Indian species of all orders of insects both in the form of preserved specimens and in the form of photographs.'],
|
| 97 |
]
|
| 98 |
scores = model.predict(pairs)
|
| 99 |
print(scores.shape)
|
|
|
|
| 101 |
|
| 102 |
# Or rank different texts based on similarity to a single text
|
| 103 |
ranks = model.rank(
|
| 104 |
+
'Khosrowabad, in the city where Pouya Tajik was born, is found in what county?',
|
| 105 |
[
|
| 106 |
+
'Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward',
|
| 107 |
+
'Frédéric Chopin. Chopin also endowed popular dance forms with a greater range of melody and expression. Chopin\'s mazurkas, while originating in the traditional Polish dance (the mazurek), differed from the traditional variety in that they were written for the concert hall rather than the dance hall; "it was Chopin who put the mazurka on the European musical map." The series of seven polonaises published in his lifetime (another nine were published posthumously), beginning with the Op. 26 pair (published 1836), set a new standard for music in the form. His waltzes were also written specifically for the salon recital rather than the ballroom and are frequently at rather faster tempos than their dance-floor equivalents.',
|
| 108 |
+
'Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward',
|
| 109 |
+
"Frédéric Chopin. Numerous recordings of Chopin's works are available. On the occasion of the composer's bicentenary, the critics of The New York Times recommended performances by the following contemporary pianists (among many others): Martha Argerich, Vladimir Ashkenazy, Emanuel Ax, Evgeny Kissin, Murray Perahia, Maurizio Pollini and Krystian Zimerman. The Warsaw Chopin Society organizes the Grand prix du disque de F. Chopin for notable Chopin recordings, held every five years.",
|
| 110 |
+
'Arignar Anna Zoological Park. The butterfly house, constructed at a cost of ₹6 million, has more than 25 host plants and landscaped habitats, such as bushes, lianas, streams, waterfall and rock - gardens, that attract many species of butterflies such as the common Mormon, crimson rose, mottled emigrant, blue tiger, evening brown and lime butterfly. A network of ponds interconnected by streams maintains humidity in the area. The park covers an area of 5 acres. The butterfly garden with an insect museum at the entrance is set up by the Tamil Nadu Agricultural University (TNAU), Coimbatore. The insect museum has been planned with an exhibit area comprising insect exhibits representing the most common Indian species of all orders of insects both in the form of preserved specimens and in the form of photographs.',
|
| 111 |
]
|
| 112 |
)
|
| 113 |
# [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]
|
|
|
|
| 149 |
| Metric | Value |
|
| 150 |
|:----------------------|:-----------|
|
| 151 |
| accuracy | 0.9648 |
|
| 152 |
+
| accuracy_threshold | 0.0927 |
|
| 153 |
| f1 | 0.9447 |
|
| 154 |
+
| f1_threshold | 0.0492 |
|
| 155 |
| precision | 0.9569 |
|
| 156 |
| recall | 0.9328 |
|
| 157 |
+
| **average_precision** | **0.9828** |
|
| 158 |
|
| 159 |
<!--
|
| 160 |
## Bias, Risks and Limitations
|
|
|
|
| 182 |
| type | string | string | float |
|
| 183 |
| details | <ul><li>min: 39 characters</li><li>mean: 76.6 characters</li><li>max: 114 characters</li></ul> | <ul><li>min: 148 characters</li><li>mean: 511.34 characters</li><li>max: 1394 characters</li></ul> | <ul><li>min: 0.0</li><li>mean: 0.32</li><li>max: 1.0</li></ul> |
|
| 184 |
* Samples:
|
| 185 |
+
| sentence_0 | sentence_1 | label |
|
| 186 |
+
|:--------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------|
|
| 187 |
+
| <code>Khosrowabad, in the city where Pouya Tajik was born, is found in what county?</code> | <code>Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward</code> | <code>1.0</code> |
|
| 188 |
+
| <code>Khosrowabad, in the city where Pouya Tajik was born, is found in what county?</code> | <code>Frédéric Chopin. Chopin also endowed popular dance forms with a greater range of melody and expression. Chopin's mazurkas, while originating in the traditional Polish dance (the mazurek), differed from the traditional variety in that they were written for the concert hall rather than the dance hall; "it was Chopin who put the mazurka on the European musical map." The series of seven polonaises published in his lifetime (another nine were published posthumously), beginning with the Op. 26 pair (published 1836), set a new standard for music in the form. His waltzes were also written specifically for the salon recital rather than the ballroom and are frequently at rather faster tempos than their dance-floor equivalents.</code> | <code>0.0</code> |
|
| 189 |
+
| <code>Margraviate of the country of the Botanical Garden of the place Josef Victor Rohon was educated is an instance of?</code> | <code>Pouya Tajik. Pouya Tajik, (born October 28, 1987 in Tehran) is a professional Iranian basketball player who currently plays for BEEM Mazandaran BC of the Iranian Super League and also for the Iranian national basketball team. He is a 6-foot-nine-inch power forward</code> | <code>0.0</code> |
|
| 190 |
* Loss: [<code>BinaryCrossEntropyLoss</code>](https://sbert.net/docs/package_reference/cross_encoder/losses.html#binarycrossentropyloss) with these parameters:
|
| 191 |
```json
|
| 192 |
{
|
|
|
|
| 324 |
### Training Logs
|
| 325 |
| Epoch | Step | validation_average_precision |
|
| 326 |
|:-----:|:----:|:----------------------------:|
|
| 327 |
+
| 1.0 | 13 | 0.9844 |
|
| 328 |
+
| 2.0 | 26 | 0.9831 |
|
| 329 |
+
| 3.0 | 39 | 0.9828 |
|
| 330 |
|
| 331 |
|
| 332 |
### Framework Versions
|
config.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
{
|
| 2 |
-
"_name_or_path": "
|
| 3 |
"architectures": [
|
| 4 |
"XLMRobertaForSequenceClassification"
|
| 5 |
],
|
|
|
|
| 1 |
{
|
| 2 |
+
"_name_or_path": "models/finetuned/musique-climb",
|
| 3 |
"architectures": [
|
| 4 |
"XLMRobertaForSequenceClassification"
|
| 5 |
],
|
eval/CrossEncoderClassificationEvaluator_validation_results.csv
CHANGED
|
@@ -2,3 +2,6 @@ epoch,steps,Accuracy,Accuracy_Threshold,F1,F1_Threshold,Precision,Recall,Average
|
|
| 2 |
1.0,13,0.94579945799458,0.00901582557708025,0.912280701754386,0.00901582557708025,0.9541284403669725,0.8739495798319328,0.9700491973968657
|
| 3 |
2.0,26,0.9512195121951219,0.033014170825481415,0.9224137931034483,0.033014170825481415,0.9469026548672567,0.8991596638655462,0.9774329046561461
|
| 4 |
3.0,39,0.964769647696477,0.04581885784864426,0.9446808510638298,0.023046962916851044,0.9568965517241379,0.9327731092436975,0.9846420260039257
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
1.0,13,0.94579945799458,0.00901582557708025,0.912280701754386,0.00901582557708025,0.9541284403669725,0.8739495798319328,0.9700491973968657
|
| 3 |
2.0,26,0.9512195121951219,0.033014170825481415,0.9224137931034483,0.033014170825481415,0.9469026548672567,0.8991596638655462,0.9774329046561461
|
| 4 |
3.0,39,0.964769647696477,0.04581885784864426,0.9446808510638298,0.023046962916851044,0.9568965517241379,0.9327731092436975,0.9846420260039257
|
| 5 |
+
1.0,13,0.964769647696477,0.09815724194049835,0.9446808510638298,0.04552947357296944,0.9568965517241379,0.9327731092436975,0.9843958194458395
|
| 6 |
+
2.0,26,0.964769647696477,0.2853778004646301,0.944206008583691,0.2853778004646301,0.9649122807017544,0.9243697478991597,0.9831305652439385
|
| 7 |
+
3.0,39,0.964769647696477,0.09269625693559647,0.9446808510638298,0.049161043018102646,0.9568965517241379,0.9327731092436975,0.9828499641745869
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2271071852
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:17af818c93fe8000999e32360bc07a8e37c1c413c3c89c95d14b4c413275a57a
|
| 3 |
size 2271071852
|
tokenizer_config.json
CHANGED
|
@@ -46,10 +46,17 @@
|
|
| 46 |
"cls_token": "<s>",
|
| 47 |
"eos_token": "</s>",
|
| 48 |
"mask_token": "<mask>",
|
|
|
|
| 49 |
"model_max_length": 1024,
|
|
|
|
| 50 |
"pad_token": "<pad>",
|
|
|
|
|
|
|
| 51 |
"sep_token": "</s>",
|
| 52 |
"sp_model_kwargs": {},
|
|
|
|
| 53 |
"tokenizer_class": "XLMRobertaTokenizer",
|
|
|
|
|
|
|
| 54 |
"unk_token": "<unk>"
|
| 55 |
}
|
|
|
|
| 46 |
"cls_token": "<s>",
|
| 47 |
"eos_token": "</s>",
|
| 48 |
"mask_token": "<mask>",
|
| 49 |
+
"max_length": 1024,
|
| 50 |
"model_max_length": 1024,
|
| 51 |
+
"pad_to_multiple_of": null,
|
| 52 |
"pad_token": "<pad>",
|
| 53 |
+
"pad_token_type_id": 0,
|
| 54 |
+
"padding_side": "right",
|
| 55 |
"sep_token": "</s>",
|
| 56 |
"sp_model_kwargs": {},
|
| 57 |
+
"stride": 0,
|
| 58 |
"tokenizer_class": "XLMRobertaTokenizer",
|
| 59 |
+
"truncation_side": "right",
|
| 60 |
+
"truncation_strategy": "longest_first",
|
| 61 |
"unk_token": "<unk>"
|
| 62 |
}
|