topk-embed-v1-xsmall-mlx

Native Apple Silicon/MLX runtime for topk-io/topk-embed-v1-xsmall, a 0.8B multimodal late-interaction retriever.

This repository keeps the source checkpoint's BF16 weights unchanged. It adds a small MLX adapter around the Qwen3.5 backbone and a fast preset used by Fenn:

  • 96 maximum visual tokens per image
  • 128-dimensional Matryoshka embeddings
  • exact-shape grouping with a maximum GPU batch size of 16
  • FP16 document/image vectors and FP32 query vectors on the host
  • MaxSim late-interaction scoring

The original checkpoint defaults to 1,280 visual tokens and 1,024 embedding dimensions. The smaller values here are inference settings, not modified or quantized weights. Increasing either setting can retain more detail at the cost of indexing speed and storage.

Requirements

  • Apple Silicon Mac
  • Python 3.11+
  • The packages in requirements.txt

Model inference runs in MLX. Image resize and patch preparation use PyTorch/torchvision on CPU.

Usage

import numpy as np
from PIL import Image

from mlx_topk_xsmall import load_topk_xsmall

model = load_topk_xsmall()

query = model.encode_queries(["flag of france"])[0]

with Image.open("page.png") as image:
    document = model.encode_images(
        [image.convert("RGB")],
        max_batch_size=16,
    )[0]

score = np.max(query @ document.T, axis=1).sum()
print(float(score))

encode_images accepts multiple images and preserves their input order. Internally it groups equal visual sequence lengths, then caps each MLX forward pass at max_batch_size. The 96-token preset is passed through load_topk_xsmall(image_token_budget=96) and can be changed by the caller.

Text document encoding is available through encode_documents. Fenn's fast multimodal mode intentionally stores pure-text rows with null semantic vectors and uses exact/keyword retrieval for those rows; that is an application-level policy, not a limitation imposed by this repository.

Attribution and license

This repository is derived from topk-io/topk-embed-v1-xsmall at revision d09d8a7a8cdd6c287f792b3c4d7b41233d46e66a. The source model and this MLX package are distributed under the Apache License 2.0. See LICENSE.

The TopK model card contains the full architecture, training, benchmark, and intended-use details.

Downloads last month
84
Safetensors
Model size
0.9B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thoddnn/topk-embed-v1-xsmall-mlx

Finetuned
(1)
this model