Sapiens2 Pose for ComfyUI (bf16 and int8 ConvRot)

Meta's Sapiens2 top-down pose estimators (0.4b, 0.8b, 1b and 5b, 308 keypoints) converted for a ComfyUI-native implementation that does not use transformers. Each size comes in two files:

  • bf16: the original weights re-keyed to the native module's names and stored in bfloat16.
  • int8 ConvRot: the transformer blocks' Linear weights quantized to per-row INT8 with ConvRot (group-wise Hadamard rotation), in ComfyUI's int8_tensorwise format. The model still computes in bfloat16.

The files are loaded by a ComfyUI node pack. A link will be added here once it is published.

Files

File Size Source (repo@revision) Recipe
sapiens2_pose_0.4b_bf16.safetensors 851,396,040 B facebook/sapiens2-pose-0.4b@73572e121602e8a6eb40eb841732e88f386d250b re-key, bf16
sapiens2_pose_0.4b_int8_convrot.safetensors 458,441,192 B same int8 ConvRot from the fp32 source, group 256
sapiens2_pose_0.8b_bf16.safetensors 1,697,464,648 B facebook/sapiens2-pose-0.8b@76f06b9854d04bb1e59e0624de39b1e8921b12aa re-key, bf16
sapiens2_pose_0.8b_int8_convrot.safetensors 886,951,960 B same int8 ConvRot from the bf16 file, group 256
sapiens2_pose_1b_bf16.safetensors 3,039,628,976 B facebook/sapiens2-pose-1b@f5fed8b97b99698d5eea1d14ff0855d0b4c3f000 re-key, bf16
sapiens2_pose_1b_int8_convrot.safetensors 1,589,457,600 B same int8 ConvRot from the bf16 file, group 256
sapiens2_pose_5b_bf16.safetensors 10,240,493,360 B facebook/sapiens2-pose-5b@ada1f29aa1fd454ca28665c700923a0101b6b24f re-key, bf16
sapiens2_pose_5b_int8_convrot.safetensors 5,184,418,256 B same int8 ConvRot from the fp32 source, group 64

Each file's safetensors metadata records its architecture (sapiens2_pose), the model config, the dtype, and the source repo and revision. The int8 files also carry _quantization_metadata and a per-layer .comfy_quant entry: {"format": "int8_tensorwise", "orig_dtype": "torch.bfloat16", "convrot": true, "convrot_groupsize": <256 or 64>, "per_row": true}.

How the files were made

bf16. The source model.safetensors (fp32, official Sapiens2 layout) is re-keyed to the native module's names and cast to bfloat16. The renames match those transformers applies when loading these repos. The fused SwiGLU w12 is kept fused. The rope_embed.periods buffer is dropped, as transformers also does, because it is recomputed from the config. Every tensor is loaded back strictly into the native module before the file is written.

int8 ConvRot. The files are quantized with convert_to_quant 1.3.4 (ctq). The command is the same for every size apart from the input and the group size:

ctq -i <source>.safetensors -o sapiens2_pose_<size>_int8_convrot.safetensors \
  --int8 --convrot --convrot-group-size <256|64> --scaling_mode row \
  --comfy_quant --save-quant-metadata --simple --low-memory \
  --exclude-layers '^patch_embed\.|^(cls|storage)_tokens|\.(ln1|ln2|q_norm|k_norm)\.|^ln_final\.|^head\.'

For an fp32 source, the input is the same re-keyed file stored in fp32 instead of bf16. After ctq, three things are set so the file computes and stores exactly like one quantized from the bf16 file:

  • every .comfy_quant entry and the _quantization_metadata get orig_dtype = torch.bfloat16;
  • every float32 tensor except the weight scales is cast to bf16;
  • every tensor was checked to have the same dtype and shape as the bf16-source variant.

Only the int8 weights, their scales, and the ConvRot-calibrated biases of the quantized layers differ. The file metadata also records quant_source_dtype: float32.

What is quantized. Every Linear in every transformer block: attn.wq, attn.wk, attn.wv, attn.proj, ffn.w12 and ffn.w3. That is 144 layers for 0.4b, 192 for 0.8b, 240 for 1b and 336 for 5b, with none skipped.

What stays bf16. The patch embedding, the class and storage tokens, all RMSNorms (ln1, ln2, q_norm, k_norm, the final norm), the layer-scale gammas, the biases and the whole heatmap head (deconvolutions, convolutions, predictor).

Why 5b uses group size 64. The 5b model's width is 2432, and 2432 is not divisible by 256. ConvRot group sizes must be powers of 4, and 64 divides both 2432 (38 × 64) and the MLP width 9728. With group 64 every Linear of 5b is quantized. The group size is read per layer from the .comfy_quant entry.

Choosing the recipe per size

For each size, several int8 candidates were made. Each was compared with this repo's bf16 file of the same size, and the one closest to it was kept. File size does not depend on the recipe.

  • Frames: 34 real person crops from one test clip: frames 0, 58, 116, 174, 232, plus 1–29, which include hands at the frame edge.
  • Keypoints and threshold: keypoints 0–69 (body, feet, hands), counted where the bf16 reference's confidence is ≥ 0.5, the drawing threshold.
  • Metrics: mean and p99 distance in frame pixels, and "flips", the keypoints that cross the 0.5 threshold in one file but not in the other.
  • Rule: lowest mean wins, with p99 as the tie-breaker. When two candidates are within 5% of each other, the simpler recipe is kept: the bf16 source and the larger group.
Size Candidate Mean px p99 px Flips
0.4b bf16 source, group 256 0.803 3.07 16
0.4b bf16 source, group 64 0.655 3.40 13
0.4b fp32 source, group 256 0.609 2.83 13 kept
0.8b bf16 source, group 256 0.477 3.09 18 kept
0.8b bf16 source, group 64 0.500 3.24 22 within 5%, simpler recipe kept
0.8b fp32 source, group 256 0.509 3.42 28
1b bf16 source, group 256 0.532 2.63 3 kept
1b bf16 source, group 64 0.596 2.37 3
1b fp32 source, group 256 0.701 2.37 4
5b bf16 source, group 64 0.526 2.05 1
5b fp32 source, group 64 0.364 1.87 1 kept

The flip counts are out of 2,380 keypoints (34 crops × 70).

Parity with the original models

These numbers compare the native module against transformers' Sapiens2ForPoseEstimation loaded from the source repo. They use frames 0, 58, 116, 174 and 232, keypoints 0–69 with the reference's confidence ≥ 0.5, and 350 keypoints in total.

Size Native fp32 vs original fp32 (mean px) Original bf16 vs original fp32 (mean px) bf16 file vs original bf16 (mean / p99 / max px, flips) int8 file vs original bf16 (mean / p99 / max px, flips)
0.4b 0.003 0.223 0.215 / 1.09 / 1.26, 2 0.333 / 1.31 / 2.8, 2
0.8b 0.006 0.226 0.187 / 0.70 / 0.86, 0 0.319 / 1.07 / 1.3, 3
1b 0.001 0.317 0.245 / 1.71 / 3.69, 0 0.324 / 1.92 / 4.6, 0
5b not measured not measured 0.165 / 0.76 / 1.13, 0 0.318 / 1.49 / 1.8, 0
  • In fp32 the native module matches the original to within a few thousandths of a pixel.
  • The bf16 files agree with the original in bf16 within the original's own bf16 rounding noise. That noise is the "Original bf16 vs original fp32" column.
  • 5b fp32 was not run, because it does not fit next to its activations on a 24 GB card.
  • Only confident keypoints are compared. On low-confidence keypoints (hidden or out of frame) the heatmap is flat, and a rounding-level difference moves its peak arbitrarily far. Those keypoints are not drawn at the 0.5 threshold.

Requirements

  • ComfyUI with int8_tensorwise + ConvRot support in comfy/ops.py (the per-layer convrot / convrot_groupsize handling) and comfy-kitchen. Tested at ComfyUI commit 79be670e with comfy-kitchen 0.2.35. The bf16 files need no quantization support.
  • The loader of the ComfyUI node pack these files are made for (link to be added). The files hold the native module's state dict, so they cannot be loaded with transformers or the original Sapiens2 code as they are.
  • Input: a 1024 × 768 person crop (top-down, one person per box), normalized with ImageNet mean/std. The output is 308 keypoint heatmaps at 256 × 192, decoded with Sapiens2's UDP pipeline.

Speed and memory

These were measured on one RTX 3090 (24 GB) with PyTorch 2.11 + CUDA 12.8 (cu128), over a 233-frame 720 × 1280 clip, one person (one crop) per frame. The time covers crop, forward and decode of all 308 keypoints.

File s/frame Peak VRAM Load
sapiens2_pose_1b_bf16 0.427 4.21 GiB 4.4 s
sapiens2_pose_1b_int8_convrot 0.615 3.26 GiB 3.3 s
sapiens2_pose_5b_bf16 0.852 10.78 GiB 7.1 s
sapiens2_pose_5b_int8_convrot 1.272 7.25 GiB 4.1 s

On this setup int8 is slower than bf16. comfy-kitchen's CUDA backend needs PyTorch built for CUDA 13.0 or newer (cu130+), so on cu128 the int8 layers run on its eager fallback. Here the gain from int8 is VRAM and disk, not speed. The 0.4b and 0.8b files were checked for parity but not timed on a full clip.

Licence

These weights are derived from Meta's Sapiens2 models (facebookresearch/sapiens2, the facebook/sapiens2-pose-* repos) and are distributed under the Sapiens2 License:

  • Included: the licence is in this repository as LICENSE.md, and it must be distributed with these files and any further derivative of them.
  • Use restrictions: by using or redistributing them you agree to its terms. Read the full text, especially the use restrictions in section 1.b.vi and the compliance requirements in sections 1.b.iii–v.
  • Attribution: Sapiens2 and the original model weights are by Meta Platforms, Inc.; this repository only re-keys and quantizes them.
  • Publications: under section 1.b.ii, research results obtained with these materials must acknowledge the use of Sapiens Materials in the publication.
  • Warranty: the files are provided as is, without any warranty; see sections 3 and 4 of the licence.
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for beycanai/sapiens2-convrot

Finetuned
(2)
this model