jarjoura's picture
Add sapiens2-pose-1b-bf16 (bf16 MLX conversion of facebook/sapiens2-pose-1b)
9c53b88 verified
|
Raw History Blame Contribute Delete
1.66 kB
metadata
license: other
license_name: sapiens2-license
license_link: https://github.com/facebookresearch/sapiens2/blob/main/LICENSE.md
pipeline_tag: keypoint-detection
library_name: mlx
base_model: facebook/sapiens2-pose-1b
tags:
  - sapiens
  - sapiens2
  - human-centric
  - pose
  - mlx
  - bf16

mlx-community/sapiens2-pose-1b-bf16

bf16 MLX conversion of facebook/sapiens2-pose-1b (Meta's Sapiens2, ICLR 2026): 308-keypoint top-down pose heatmaps. Converted with mlx-vlm 0.7.0; the original float32 checkpoint is 2x this size.

What is in model.safetensors (3.04 GB):

  • every parameter in bfloat16 (the reference runs inference in bf16 mixed precision);
  • the q/k/v projections merged into one wqkv tensor per block, as the mlx-vlm Sapiens2 model expects.

Refer to the original model card for the model description, intended use and license.

Use with mlx-vlm

pip install -U mlx-vlm
from mlx_vlm import load
from mlx_vlm.models.sapiens2.generate import Sapiens2Predictor, read_image

model, _ = load("mlx-community/sapiens2-pose-1b-bf16")
predictor = Sapiens2Predictor(model)
# boxes: (N, 4) xyxy person boxes from a detector; defaults to the full image.
output = predictor.infer(read_image("image.jpg"), boxes=boxes, flip_test=False)
keypoints, scores = output["keypoints"], output["scores"]  # (N, 308, 2), (N, 308)

Outputs are numpy arrays at the input resolution (dense tasks) or in source-image pixel coordinates (pose). See the mlx-vlm Sapiens2 README for preprocessing details and the per-task output keys.