Phonon-1-Big / README.md
FermionResearch's picture
Run it: CPU
9a3805c verified
|
Raw
History Blame Contribute Delete
1.62 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-ASR-0.6B
base_model_relation: quantized
language:
- en
library_name: mlx
pipeline_tag: automatic-speech-recognition
tags:
- mlx
- apple-silicon
- speech-to-text
- asr
- stt
- low-bit
- ternary
- quantization-aware-training
- on-device
- streaming
metrics:
- wer
---
# Phonon-1 Big
This is the largest build of the Phonon-1 family, an open speech recognition
model for English that downloads in 581 MB.
## Benchmarks
| Benchmark | Phonon-1 Big (581 MB) |
|---|---:|
| LibriSpeech test-clean | 2.667 |
| LibriSpeech test-other | 5.722 |
| TED-LIUM | 3.400 |
| SPGISpeech | 4.156 |
| VoxPopuli | 8.369 |
| GigaSpeech | 11.291 |
| Earnings-22 | 12.417 |
| AMI | 12.812 |
| Macro (eight benchmarks) | 7.604 |
Word error rate, lower is better. Measured by us — full test sets, Whisper
English text normalizer, greedy decoding.
## Run it
```bash
pip install fermion-research
fermion transcribe recording.wav --model FermionResearch/Phonon-1-Big
```
Or serve an OpenAI-compatible endpoint:
```bash
fermion serve --model FermionResearch/Phonon-1-Big
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F "file=@recording.wav" \
-F "model=FermionResearch/Phonon-1-Big"
```
The same weights run on a Mac (via MLX), on an NVIDIA GPU, or on a plain CPU;
the runtimes and Docker images are in the [GitHub repo](https://github.com/fermionresearch/phonon).
## License
**Apache License 2.0** for the weights and the [command line](https://pypi.org/project/fermion-research/). Base model: [`Qwen/Qwen3-ASR-0.6B`](https://huggingface.co/Qwen/Qwen3-ASR-0.6B), Apache-2.0.