File size: 6,463 Bytes
4ca6408 b94bf59 4ca6408 b94bf59 4ca6408 b94bf59 4ca6408 b94bf59 4ca6408 b94bf59 4ca6408 b94bf59 4ca6408 b94bf59 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | ---
license: apache-2.0
language:
- en
pipeline_tag: feature-extraction
tags:
- audio-retrieval
- embedding
- custom_code
- multimodal
base_model: mispeech/midashenglm-7b-0804-fp32
---
<h1 align="center">ALM2Vec-FT</h1>
<p align="center">
<a href="https://arxiv.org/abs/xxxx.xxxxx"><img src="https://img.shields.io/badge/Paper-arXiv-b31b1b?logo=arxiv&logoColor=white" alt="Paper"></a>
<a href="https://caml-labs.github.io/ALM2Vec"><img src="https://img.shields.io/badge/Project-Page-1f6feb?logo=googlechrome&logoColor=white" alt="Project Page"></a>
<a href="https://github.com/caml-labs/ALM2Vec"><img src="https://img.shields.io/badge/Code-GitHub-181717?logo=github&logoColor=white" alt="GitHub"></a>
</p>
**ALM2Vec** is a universal audio embedding model for retrieval, derived from a pretrained large audio–language model (LALM). Instead of being optimized only for audio–caption matching like conventional contrastive dual-encoders, it transfers the audio understanding, instruction-following, and reasoning abilities of LALMs into a single unified embedding space that works across audio domains, task types, and user intents.
Its key feature is **instruction-aware retrieval**: a natural-language instruction guides the embedding, so the *same* audio can be encoded differently for different needs. This supports:
- **Instruction-aware retrieval** — focus the embedding on a specific aspect of the audio.
- **Text ↔ audio retrieval** — bidirectional matching between audio and text.
- **Audio question answering** — match an audio query plus a question against candidate answers.
ALM2Vec achieves competitive results on standard audio and speech retrieval benchmarks while adding these controllable retrieval capabilities. See the [project page](https://caml-labs.github.io/ALM2Vec/) for interactive demos.
This repository hosts the **finetune** checkpoint, built on [MiDashengLM](https://huggingface.co/mispeech/midashenglm-7b-0804-fp32).
Requirements: `transformers>=4.52`, `torch`, `safetensors`, and `torchaudio` for non-WAV audio. Requires a GPU (~31GB weights) and `trust_remote_code=True`.
## Example
```python
import torch
from transformers import AutoModel, AutoTokenizer
repo_id = "cara-ai/ALM2Vec-FT"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
repo_id, trust_remote_code=True, torch_dtype=torch.float32
).cuda().eval()
model.set_tokenizer(tokenizer)
query_text = ["q1", "q2"]
query_audio = ["/path/to/audio1.wav", "https://example.com/audio2.wav"]
doc_text = ["d1", "d2"]
query_emb = model.encode_query(text=query_text, audio=query_audio, task="query")
doc_emb = model.encode_document(text=doc_text, task="document")
similarity = query_emb @ doc_emb.T
print(similarity)
```
## Results
**ALM2Vec-FT** is the checkpoint hosted in this repository; **ALM2Vec-PT** is the pretrain variant. In every table, **bold** marks the best score and <u>underline</u> the second best.
### Text–audio retrieval — AudioCaps
| Method | T→A R@1 | T→A R@5 | T→A R@10 | A→T R@1 | A→T R@5 | A→T R@10 |
| --- | --- | --- | --- | --- | --- | --- |
| LAION-CLAP | 36.1 | 71.8 | 83.9 | 46.8 | <u>82.9</u> | <u>90.7</u> |
| MS-CLAP | 15.4 | 47.2 | 64.5 | 32.0 | 66.0 | 79.2 |
| WavCaps-CLAP-PT | 39.7 | 74.5 | 86.1 | 51.7 | 82.3 | 90.6 |
| WavCaps-CLAP-FT | <u>42.2</u> | <u>76.5</u> | <u>87.1</u> | <u>54.6</u> | **85.2** | **92.4** |
| JINA-Embed.-v5 | 20.4 | 50.3 | 64.4 | 23.1 | 52.7 | 67.2 |
| **ALM2Vec-PT** | 40.0 | 74.5 | 85.9 | 43.8 | 74.3 | 86.5 |
| **ALM2Vec-FT** | **43.2** | **78.0** | **87.8** | **55.5** | 80.0 | 88.2 |
### Text–audio retrieval — Clotho
| Method | T→A R@1 | T→A R@5 | T→A R@10 | A→T R@1 | A→T R@5 | A→T R@10 |
| --- | --- | --- | --- | --- | --- | --- |
| LAION-CLAP | 16.1 | 38.3 | 51.1 | 22.7 | 48.5 | 60.8 |
| MS-CLAP | 15.6 | 38.9 | 51.4 | 22.1 | 48.9 | 62.0 |
| WavCaps-CLAP-PT | 19.5 | 45.2 | 58.2 | 23.4 | 50.9 | 63.4 |
| WavCaps-CLAP-FT | <u>19.7</u> | <u>45.7</u> | <u>59.4</u> | <u>26.9</u> | <u>52.6</u> | <u>64.9</u> |
| JINA-Embed.-v5 | 9.2 | 23.9 | 35.0 | 10.5 | 24.7 | 34.3 |
| **ALM2Vec-PT** | 19.2 | 43.4 | 55.7 | 17.9 | 39.4 | 52.2 |
| **ALM2Vec-FT** | **24.8** | **52.9** | **65.8** | **27.9** | **52.7** | **66.3** |
### Speech retrieval — LibriSQA
| Method | T→S R@1 | T→S R@5 | T→S R@10 | S→T R@1 | S→T R@5 | S→T R@10 |
| -------------- | --------------- | -------- | -------- | --------------- | -------- | -------- |
| LAION-CLAP † | 0.0 | 0.1 | 0.8 | 0.1 | 0.2 | 0.6 |
| Whisper+BGE | 83.7 | 93.3 | 94.9 | 85.2 | 93.4 | 95.3 |
| CLSR | **85.0** | <u>93.4</u> | <u>95.0</u> | <u>85.5</u> | <u>94.0</u> | <u>95.6</u> |
| **ALM2Vec-PT** | 43.7 | 64.5 | 72.8 | 11.2 | 24.9 | 34.1 |
| **ALM2Vec-FT** | <u>84.7</u> | **94.1** | **95.8** | **86.0** | **95.2** | **97.2** |
### Audio understanding — MMAU-mini (accuracy)
| Method | Overall | Music | Sound | Speech |
| ------------------ | -------- | -------- | -------- | -------- |
| GPT-4o Audio ‡ | 60.8 | 63.2 | 64.6 | 56.3 |
| Gemini 2.5 Pro ‡ | <u>71.6</u> | <u>75.1</u> | 71.5 | 68.3 |
| Qwen2.5-Omni ‡ | 71.5 | 65.9 | <u>78.1</u> | <u>70.6</u> |
| Audio Flamingo 3 ‡ | **73.1** | **76.9** | 66.1 | **73.9** |
| **ALM2Vec-PT** | 66.3 | 62.3 | **78.7** | 58.0 |
| **ALM2Vec-FT** | 63.0 | 61.7 | 74.8 | 52.6 |
† LAION-CLAP is not trained for speech and effectively fails on LibriSQA; shown for reference.
‡ Generative large audio–language models, listed as reference upper bounds rather than directly comparable retrieval baselines.
## Citation
If you find this work useful, please consider citing:
```
@article{ALM2Vec2026,
title={ALM2Vec: Learning Audio Embeddings for Universal
Audio Retrieval with Large Audio-Language Models},
author={TBD},
journal={arXiv preprint arXiv:TBD},
year={2026}
}
```
## Acknowledgement
ALM2Vec is built on [MiDashengLM](https://github.com/xiaomi-research/dasheng-lm) and further trained for universal audio retrieval. We thank MiDashengLM and its underlying [Dasheng](https://github.com/RicherMans/Dasheng) audio encoder for their open-source contributions. |