Instructions to use vllm-sr/Vela-1.0-Encoder-307M-Embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use vllm-sr/Vela-1.0-Encoder-307M-Embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("vllm-sr/Vela-1.0-Encoder-307M-Embedding") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Download AMD_RUNTIME.md from vllm-sr/Vela-1.0-Encoder-307M-Embedding: direct link, hf CLI and curl.
- Browser
- Download file 2.94 kB
-
https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Embedding/resolve/main/AMD_RUNTIME.md
- Command line
-
hf download hf://vllm-sr/Vela-1.0-Encoder-307M-Embedding/AMD_RUNTIME.md
-
curl -L -o AMD_RUNTIME.md https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Embedding/resolve/main/AMD_RUNTIME.md
Vela inference on AMD ROCm
The standard Sentence Transformers embedding loader does not automatically install the optional helper described here. Enable it explicitly before the first inference when using the affected runtime. The reranker custom loader can install the same helper directly.
Qualified runtime
The issue was reproduced with PyTorch 2.8.0 + ROCm 6.4, Transformers 4.57.6, Sentence Transformers 5.1.2, and native ModernBERT SDPA attention in FP32. The efficient attention backend produced different outputs from MATH when its post-RoPE Q/K and packed-projection V inputs were non-contiguous. Making those three inputs contiguous restored agreement without changing weights, pooling, masking, or scoring heads.
This qualification covers the published Vela embedding checkpoint at
75c25cdc3072e85872c82a77a64a749bcef5f5e5. It does not establish behavior for
other package versions, backends, or exported ONNX models. The existing
embedding training and evaluation runs used the MATH backend.
Explicit helper
Review modernbert_sdpa_layout.py, then place it
beside your application. For the qualified file, its SHA256 is
0fb3a22db93ad76e30dfbfb3011de442149d97c565e139a3a55738287a1fbbc8.
The file must come from the same version as these instructions.
import torch
from sentence_transformers import SentenceTransformer
from modernbert_sdpa_layout import install_rocm_sdpa_layout_guard
model = SentenceTransformer(
"vllm-sr/Vela-1.0-Encoder-307M-Embedding",
revision="75c25cdc3072e85872c82a77a64a749bcef5f5e5",
device="cuda",
trust_remote_code=False,
model_kwargs={"torch_dtype": torch.float32, "attn_implementation": "sdpa"},
config_kwargs={"reference_compile": False},
)
install_rocm_sdpa_layout_guard(model[0].auto_model)
vectors = model.encode(
["How do I reset my password?", "I forgot my account password."],
normalize_embeddings=True,
)
Installation changes only this loaded model's attention methods. CPU,
non-ROCm devices, and forced MATH calls keep their original implementation.
The helper is idempotent and does not add parameters or change checkpoint
keys. Install it again after loading another model, including a model saved
with save_pretrained; Python method replacements are not serialized.
MATH alternative
Without the helper, run inference inside an explicit MATH context:
from torch.nn.attention import SDPBackend, sdpa_kernel
with sdpa_kernel(SDPBackend.MATH):
vectors = model.encode(
["How do I reset my password?", "I forgot my account password."],
normalize_embeddings=True,
)
This preserves the standard loader API but can require more memory and time, especially for long inputs. Do not interpret an unqualified attention backend as a model-quality improvement or compare scores across different arithmetic paths without first checking numerical agreement.