Instructions to use thoddnn/topk-embed-v1-xsmall-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use thoddnn/topk-embed-v1-xsmall-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download thoddnn/topk-embed-v1-xsmall-mlx --local-dir topk-embed-v1-xsmall-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
topk-embed-v1-xsmall-mlx
Native Apple Silicon/MLX runtime for topk-io/topk-embed-v1-xsmall, a 0.8B multimodal late-interaction retriever.
This repository keeps the source checkpoint's BF16 weights unchanged. It adds a small MLX adapter around the Qwen3.5 backbone and a fast preset used by Fenn:
- 96 maximum visual tokens per image
- 128-dimensional Matryoshka embeddings
- exact-shape grouping with a maximum GPU batch size of 16
- FP16 document/image vectors and FP32 query vectors on the host
- MaxSim late-interaction scoring
The original checkpoint defaults to 1,280 visual tokens and 1,024 embedding dimensions. The smaller values here are inference settings, not modified or quantized weights. Increasing either setting can retain more detail at the cost of indexing speed and storage.
Requirements
- Apple Silicon Mac
- Python 3.11+
- The packages in
requirements.txt
Model inference runs in MLX. Image resize and patch preparation use PyTorch/torchvision on CPU.
Usage
import numpy as np
from PIL import Image
from mlx_topk_xsmall import load_topk_xsmall
model = load_topk_xsmall()
query = model.encode_queries(["flag of france"])[0]
with Image.open("page.png") as image:
document = model.encode_images(
[image.convert("RGB")],
max_batch_size=16,
)[0]
score = np.max(query @ document.T, axis=1).sum()
print(float(score))
encode_images accepts multiple images and preserves their input order. Internally it groups equal visual sequence lengths, then caps each MLX forward pass at max_batch_size. The 96-token preset is passed through load_topk_xsmall(image_token_budget=96) and can be changed by the caller.
Text document encoding is available through encode_documents. Fenn's fast multimodal mode intentionally stores pure-text rows with null semantic vectors and uses exact/keyword retrieval for those rows; that is an application-level policy, not a limitation imposed by this repository.
Attribution and license
This repository is derived from topk-io/topk-embed-v1-xsmall at revision d09d8a7a8cdd6c287f792b3c4d7b41233d46e66a. The source model and this MLX package are distributed under the Apache License 2.0. See LICENSE.
The TopK model card contains the full architecture, training, benchmark, and intended-use details.
- Downloads last month
- 84
Quantized
Model tree for thoddnn/topk-embed-v1-xsmall-mlx
Base model
Qwen/Qwen3.5-0.8B-Base