--- license: other license_name: adobe-research-license license_link: https://github.com/adobe-research/speaker-identification/blob/main/LICENSE.md library_name: pytorch base_model: FacebookAI/roberta-large tags: - speaker-identification - quantization - int8 - roberta - mediasum --- # Joint SpeakerID INT8 (weight-only) INT8 weight-only quantisation of the Joint Speaker Identifier from [adobe-research/speaker-identification](https://github.com/adobe-research/speaker-identification) ([Interspeech 2024](https://arxiv.org/abs/2407.12094)). This checkpoint is a **quantised child** of Adobe Research’s original Joint Speaker Identifier (FP32). The architecture and trained weights are Adobe’s; only the Linear weights were packed to INT8 (group size 64). No additional training. | Metric | Value | |---|---| | Precision | **78.87** | | Δ vs FP32 parent | **0.00** | | F1 | 63.28 | | Accuracy | 67.60 | | Throughput | 18.34 examples/s (Apple M3 Pro, MPS, batch 2) | | In-memory size | 472 MB (FP32 parent: 1633 MB) | Same precision, F1, and accuracy as the FP32 Joint parent at about 3.5× smaller weights. ## Parent model | | | |---|---| | Parent | Joint Speaker Identifier (FP32), Adobe Research | | Original weights | `logs/mediasum-joint/best-model.mdl` in [adobe-research/speaker-identification](https://github.com/adobe-research/speaker-identification) | | Paper | [Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models](https://arxiv.org/abs/2407.12094) (Interspeech 2024) | | Backbone | [FacebookAI/roberta-large](https://huggingface.co/FacebookAI/roberta-large) | | Relation | weight-only INT8 quantisation of the Adobe Joint checkpoint | ## Files - `model.pt` — quantised bundle (`scheme=int8_weight`, group size 64) - `config.json` — metrics and load metadata ## Load ```python from huggingface_hub import hf_hub_download path = hf_hub_download("hmarchant/speaker-id-joint-int8", "model.pt") ``` ## License Derived from Adobe Research Speaker Identification. The [Adobe Research License](https://github.com/adobe-research/speaker-identification/blob/main/LICENSE.md) allows **non-commercial research use only**.