Xunzhuo's picture
Update Hugging Face org references: llm-semantic-router → vllm-sr
290318f verified
|
Raw History Blame Contribute Delete
2.94 kB

Vela inference on AMD ROCm

The standard Sentence Transformers embedding loader does not automatically install the optional helper described here. Enable it explicitly before the first inference when using the affected runtime. The reranker custom loader can install the same helper directly.

Qualified runtime

The issue was reproduced with PyTorch 2.8.0 + ROCm 6.4, Transformers 4.57.6, Sentence Transformers 5.1.2, and native ModernBERT SDPA attention in FP32. The efficient attention backend produced different outputs from MATH when its post-RoPE Q/K and packed-projection V inputs were non-contiguous. Making those three inputs contiguous restored agreement without changing weights, pooling, masking, or scoring heads.

This qualification covers the published Vela embedding checkpoint at 75c25cdc3072e85872c82a77a64a749bcef5f5e5. It does not establish behavior for other package versions, backends, or exported ONNX models. The existing embedding training and evaluation runs used the MATH backend.

Explicit helper

Review modernbert_sdpa_layout.py, then place it beside your application. For the qualified file, its SHA256 is 0fb3a22db93ad76e30dfbfb3011de442149d97c565e139a3a55738287a1fbbc8. The file must come from the same version as these instructions.

import torch
from sentence_transformers import SentenceTransformer

from modernbert_sdpa_layout import install_rocm_sdpa_layout_guard

model = SentenceTransformer(
    "vllm-sr/Vela-1.0-Encoder-307M-Embedding",
    revision="75c25cdc3072e85872c82a77a64a749bcef5f5e5",
    device="cuda",
    trust_remote_code=False,
    model_kwargs={"torch_dtype": torch.float32, "attn_implementation": "sdpa"},
    config_kwargs={"reference_compile": False},
)
install_rocm_sdpa_layout_guard(model[0].auto_model)
vectors = model.encode(
    ["How do I reset my password?", "I forgot my account password."],
    normalize_embeddings=True,
)

Installation changes only this loaded model's attention methods. CPU, non-ROCm devices, and forced MATH calls keep their original implementation. The helper is idempotent and does not add parameters or change checkpoint keys. Install it again after loading another model, including a model saved with save_pretrained; Python method replacements are not serialized.

MATH alternative

Without the helper, run inference inside an explicit MATH context:

from torch.nn.attention import SDPBackend, sdpa_kernel

with sdpa_kernel(SDPBackend.MATH):
    vectors = model.encode(
        ["How do I reset my password?", "I forgot my account password."],
        normalize_embeddings=True,
    )

This preserves the standard loader API but can require more memory and time, especially for long inputs. Do not interpret an unqualified attention backend as a model-quality improvement or compare scores across different arithmetic paths without first checking numerical agreement.