File size: 2,314 Bytes
f80556d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d52dd40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f80556d
 
 
 
 
 
 
 
 
 
 
 
 
15a3bec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
license: mit
language:
- fr
base_model:
- almanach/camembert-base
pipeline_tag: text-classification
library_name: transformers
tags:
- discourse relation
- discourse connective
- camembert
- nlp
model-index:
- name: Relex
  results:
  - task:
      type: text-classification
    metrics:
    - name: Label Names
      type: list
      value: 
        - alternation
        - background
        - commentary
        - concession
        - condition
        - consequence
        - continuation
        - contrast
        - detachment
        - evidence
        - explanation
        - explanation*
        - flashback
        - goal
        - narration
        - parallel
        - result
        - result*
        - summary
    - name: macro-F1
      type: f1
      value: 0.59
    - name: Accuracy
      type: accuracy
      value: 0.63
    - name: Precision
      type: precision
      value: 0.62
    - name: Recall
      type: recall
      value: 0.62

---
# Model description

*Relex* is a fine-tuned CamemBERT model trained to classify the relation expressed by a connective in context. Given a connective tagged by the tokens [MARKER] and [/MARKER], Relex predicts the relation of this connective.

- *Training data*: French newspapers and Wikiconflit comments, automatically annotated in connectives

- *Special tokens*: Connectives are wrapped between [MARKER] and [/MARKER] tokens in the training data. These tags signal to the model which word it should focus its attention on for the relation mapping.

- *Context Window*: The special tokens must appear within the first 256 tokens of the input. Because these signals are the anchor for the classification, ensuring they are not truncated is crucial for accurate predictions.

- *Predictions*: Relex predicts among 19 discourse relations (SDRT) .

- *Example*:
- *Input*: [MARKER] Peu avant de [/MARKER] mourir, Mio a promis à son mari qu'elle reviendrait à la saison des pluies.
- *Prediction*: Narration

# Usage

You can use this model directly with a Hugging Face pipeline:

```python
from transformers import pipeline

pipe = pipeline("text-classification", model="FatouSow/Relex")

text ="[MARKER] Peu avant de [/MARKER] mourir, Mio a promis à son mari qu'elle reviendrait à la saison des pluies."

result = pipe(text)
print(result)
```