hmarchant commited on
Commit
9ec7b17
·
verified ·
1 Parent(s): 3677e3a

Attribute as quantised child of Adobe Joint SpeakerID; drop GitHub fork links

Browse files
Files changed (1) hide show
  1. README.md +18 -10
README.md CHANGED
@@ -3,6 +3,7 @@ license: other
3
  license_name: adobe-research-license
4
  license_link: https://github.com/adobe-research/speaker-identification/blob/main/LICENSE.md
5
  library_name: pytorch
 
6
  tags:
7
  - speaker-identification
8
  - quantization
@@ -17,20 +18,33 @@ INT8 weight-only quantisation of the Joint Speaker Identifier from
17
  [adobe-research/speaker-identification](https://github.com/adobe-research/speaker-identification)
18
  ([Interspeech 2024](https://arxiv.org/abs/2407.12094)).
19
 
 
 
 
 
 
20
  **Joint INT8 — 78.87 / 0.00 / 63.28 / 67.60 / 18.34 / 472 MB**
21
 
22
  | Metric | Value |
23
  |---|---|
24
  | Precision | **78.87** |
25
- | Δ vs FP32 | **0.00** |
26
  | F1 | 63.28 |
27
  | Accuracy | 67.60 |
28
  | Throughput | 18.34 examples/s (Apple M3 Pro, MPS, batch 2) |
29
- | In-memory size | 472 MB (FP32: 1633 MB) |
 
 
30
 
31
- Same precision, F1, and accuracy as the FP32 Joint baseline at about 3.5× smaller weights.
32
 
33
- Code and eval tables: [HugoMarchant/speaker-identification-quantisation](https://github.com/HugoMarchant/speaker-identification-quantisation).
 
 
 
 
 
 
34
 
35
  ## Files
36
 
@@ -41,16 +55,10 @@ Code and eval tables: [HugoMarchant/speaker-identification-quantisation](https:/
41
 
42
  ```python
43
  from huggingface_hub import hf_hub_download
44
- from device_utils import apply_cuda_shims
45
- apply_cuda_shims()
46
- from model_io import load_quantized_bundle
47
 
48
  path = hf_hub_download("hmarchant/speaker-id-joint-int8", "model.pt")
49
- model, config, tokenizer, qcfg, info, device = load_quantized_bundle(path)
50
  ```
51
 
52
- Requires the companion GitHub repo (custom `JointSpeakerIdentifier` + `WeightOnlyQuantizedLinear`).
53
-
54
  ## License
55
 
56
  Derived from Adobe Research Speaker Identification. The
 
3
  license_name: adobe-research-license
4
  license_link: https://github.com/adobe-research/speaker-identification/blob/main/LICENSE.md
5
  library_name: pytorch
6
+ base_model: FacebookAI/roberta-large
7
  tags:
8
  - speaker-identification
9
  - quantization
 
18
  [adobe-research/speaker-identification](https://github.com/adobe-research/speaker-identification)
19
  ([Interspeech 2024](https://arxiv.org/abs/2407.12094)).
20
 
21
+ This checkpoint is a **quantised child** of Adobe Research’s original Joint
22
+ Speaker Identifier (FP32). The architecture and trained weights are Adobe’s;
23
+ only the Linear weights were packed to INT8 (group size 64). No additional
24
+ training.
25
+
26
  **Joint INT8 — 78.87 / 0.00 / 63.28 / 67.60 / 18.34 / 472 MB**
27
 
28
  | Metric | Value |
29
  |---|---|
30
  | Precision | **78.87** |
31
+ | Δ vs FP32 parent | **0.00** |
32
  | F1 | 63.28 |
33
  | Accuracy | 67.60 |
34
  | Throughput | 18.34 examples/s (Apple M3 Pro, MPS, batch 2) |
35
+ | In-memory size | 472 MB (FP32 parent: 1633 MB) |
36
+
37
+ Same precision, F1, and accuracy as the FP32 Joint parent at about 3.5× smaller weights.
38
 
39
+ ## Parent model
40
 
41
+ | | |
42
+ |---|---|
43
+ | Parent | Joint Speaker Identifier (FP32), Adobe Research |
44
+ | Original weights | `logs/mediasum-joint/best-model.mdl` in [adobe-research/speaker-identification](https://github.com/adobe-research/speaker-identification) |
45
+ | Paper | [Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models](https://arxiv.org/abs/2407.12094) (Interspeech 2024) |
46
+ | Backbone | [FacebookAI/roberta-large](https://huggingface.co/FacebookAI/roberta-large) |
47
+ | Relation | weight-only INT8 quantisation of the Adobe Joint checkpoint |
48
 
49
  ## Files
50
 
 
55
 
56
  ```python
57
  from huggingface_hub import hf_hub_download
 
 
 
58
 
59
  path = hf_hub_download("hmarchant/speaker-id-joint-int8", "model.pt")
 
60
  ```
61
 
 
 
62
  ## License
63
 
64
  Derived from Adobe Research Speaker Identification. The