Sapiens2 Pose for ComfyUI (bf16 and int8 ConvRot)
Meta's Sapiens2 top-down pose estimators (0.4b, 0.8b, 1b and 5b, 308 keypoints) converted for a
ComfyUI-native implementation that does not use transformers. Each size comes in two files:
- bf16: the original weights re-keyed to the native module's names and stored in bfloat16.
- int8 ConvRot: the transformer blocks' Linear weights quantized to per-row INT8 with ConvRot
(group-wise Hadamard rotation), in ComfyUI's
int8_tensorwiseformat. The model still computes in bfloat16.
The files are loaded by a ComfyUI node pack. A link will be added here once it is published.
Files
| File | Size | Source (repo@revision) | Recipe |
|---|---|---|---|
sapiens2_pose_0.4b_bf16.safetensors |
851,396,040 B | facebook/sapiens2-pose-0.4b@73572e121602e8a6eb40eb841732e88f386d250b |
re-key, bf16 |
sapiens2_pose_0.4b_int8_convrot.safetensors |
458,441,192 B | same | int8 ConvRot from the fp32 source, group 256 |
sapiens2_pose_0.8b_bf16.safetensors |
1,697,464,648 B | facebook/sapiens2-pose-0.8b@76f06b9854d04bb1e59e0624de39b1e8921b12aa |
re-key, bf16 |
sapiens2_pose_0.8b_int8_convrot.safetensors |
886,951,960 B | same | int8 ConvRot from the bf16 file, group 256 |
sapiens2_pose_1b_bf16.safetensors |
3,039,628,976 B | facebook/sapiens2-pose-1b@f5fed8b97b99698d5eea1d14ff0855d0b4c3f000 |
re-key, bf16 |
sapiens2_pose_1b_int8_convrot.safetensors |
1,589,457,600 B | same | int8 ConvRot from the bf16 file, group 256 |
sapiens2_pose_5b_bf16.safetensors |
10,240,493,360 B | facebook/sapiens2-pose-5b@ada1f29aa1fd454ca28665c700923a0101b6b24f |
re-key, bf16 |
sapiens2_pose_5b_int8_convrot.safetensors |
5,184,418,256 B | same | int8 ConvRot from the fp32 source, group 64 |
Each file's safetensors metadata records its architecture (sapiens2_pose), the model config, the
dtype, and the source repo and revision. The int8 files also carry _quantization_metadata and a
per-layer .comfy_quant entry: {"format": "int8_tensorwise", "orig_dtype": "torch.bfloat16", "convrot": true, "convrot_groupsize": <256 or 64>, "per_row": true}.
How the files were made
bf16. The source model.safetensors (fp32, official Sapiens2 layout) is re-keyed to the native
module's names and cast to bfloat16. The renames match those transformers applies when loading these
repos. The fused SwiGLU w12 is kept fused. The rope_embed.periods buffer is dropped, as
transformers also does, because it is recomputed from the config. Every tensor is loaded back
strictly into the native module before the file is written.
int8 ConvRot. The files are quantized with
convert_to_quant 1.3.4 (ctq). The command is the
same for every size apart from the input and the group size:
ctq -i <source>.safetensors -o sapiens2_pose_<size>_int8_convrot.safetensors \
--int8 --convrot --convrot-group-size <256|64> --scaling_mode row \
--comfy_quant --save-quant-metadata --simple --low-memory \
--exclude-layers '^patch_embed\.|^(cls|storage)_tokens|\.(ln1|ln2|q_norm|k_norm)\.|^ln_final\.|^head\.'
For an fp32 source, the input is the same re-keyed file stored in fp32 instead of bf16. After ctq,
three things are set so the file computes and stores exactly like one quantized from the bf16 file:
- every
.comfy_quantentry and the_quantization_metadatagetorig_dtype=torch.bfloat16; - every float32 tensor except the weight scales is cast to bf16;
- every tensor was checked to have the same dtype and shape as the bf16-source variant.
Only the int8 weights, their scales, and the ConvRot-calibrated biases of the quantized layers differ.
The file metadata also records quant_source_dtype: float32.
What is quantized. Every Linear in every transformer block: attn.wq, attn.wk, attn.wv,
attn.proj, ffn.w12 and ffn.w3. That is 144 layers for 0.4b, 192 for 0.8b, 240 for 1b and 336 for
5b, with none skipped.
What stays bf16. The patch embedding, the class and storage tokens, all RMSNorms (ln1, ln2,
q_norm, k_norm, the final norm), the layer-scale gammas, the biases and the whole heatmap head
(deconvolutions, convolutions, predictor).
Why 5b uses group size 64. The 5b model's width is 2432, and 2432 is not divisible by 256. ConvRot
group sizes must be powers of 4, and 64 divides both 2432 (38 × 64) and the MLP width 9728. With group
64 every Linear of 5b is quantized. The group size is read per layer from the .comfy_quant entry.
Choosing the recipe per size
For each size, several int8 candidates were made. Each was compared with this repo's bf16 file of the same size, and the one closest to it was kept. File size does not depend on the recipe.
- Frames: 34 real person crops from one test clip: frames 0, 58, 116, 174, 232, plus 1–29, which include hands at the frame edge.
- Keypoints and threshold: keypoints 0–69 (body, feet, hands), counted where the bf16 reference's confidence is ≥ 0.5, the drawing threshold.
- Metrics: mean and p99 distance in frame pixels, and "flips", the keypoints that cross the 0.5 threshold in one file but not in the other.
- Rule: lowest mean wins, with p99 as the tie-breaker. When two candidates are within 5% of each other, the simpler recipe is kept: the bf16 source and the larger group.
| Size | Candidate | Mean px | p99 px | Flips | |
|---|---|---|---|---|---|
| 0.4b | bf16 source, group 256 | 0.803 | 3.07 | 16 | |
| 0.4b | bf16 source, group 64 | 0.655 | 3.40 | 13 | |
| 0.4b | fp32 source, group 256 | 0.609 | 2.83 | 13 | kept |
| 0.8b | bf16 source, group 256 | 0.477 | 3.09 | 18 | kept |
| 0.8b | bf16 source, group 64 | 0.500 | 3.24 | 22 | within 5%, simpler recipe kept |
| 0.8b | fp32 source, group 256 | 0.509 | 3.42 | 28 | |
| 1b | bf16 source, group 256 | 0.532 | 2.63 | 3 | kept |
| 1b | bf16 source, group 64 | 0.596 | 2.37 | 3 | |
| 1b | fp32 source, group 256 | 0.701 | 2.37 | 4 | |
| 5b | bf16 source, group 64 | 0.526 | 2.05 | 1 | |
| 5b | fp32 source, group 64 | 0.364 | 1.87 | 1 | kept |
The flip counts are out of 2,380 keypoints (34 crops × 70).
Parity with the original models
These numbers compare the native module against transformers' Sapiens2ForPoseEstimation loaded from
the source repo. They use frames 0, 58, 116, 174 and 232, keypoints 0–69 with the reference's confidence
≥ 0.5, and 350 keypoints in total.
| Size | Native fp32 vs original fp32 (mean px) | Original bf16 vs original fp32 (mean px) | bf16 file vs original bf16 (mean / p99 / max px, flips) | int8 file vs original bf16 (mean / p99 / max px, flips) |
|---|---|---|---|---|
| 0.4b | 0.003 | 0.223 | 0.215 / 1.09 / 1.26, 2 | 0.333 / 1.31 / 2.8, 2 |
| 0.8b | 0.006 | 0.226 | 0.187 / 0.70 / 0.86, 0 | 0.319 / 1.07 / 1.3, 3 |
| 1b | 0.001 | 0.317 | 0.245 / 1.71 / 3.69, 0 | 0.324 / 1.92 / 4.6, 0 |
| 5b | not measured | not measured | 0.165 / 0.76 / 1.13, 0 | 0.318 / 1.49 / 1.8, 0 |
- In fp32 the native module matches the original to within a few thousandths of a pixel.
- The bf16 files agree with the original in bf16 within the original's own bf16 rounding noise. That noise is the "Original bf16 vs original fp32" column.
- 5b fp32 was not run, because it does not fit next to its activations on a 24 GB card.
- Only confident keypoints are compared. On low-confidence keypoints (hidden or out of frame) the heatmap is flat, and a rounding-level difference moves its peak arbitrarily far. Those keypoints are not drawn at the 0.5 threshold.
Requirements
- ComfyUI with
int8_tensorwise+ ConvRot support incomfy/ops.py(the per-layerconvrot/convrot_groupsizehandling) andcomfy-kitchen. Tested at ComfyUI commit79be670ewith comfy-kitchen 0.2.35. The bf16 files need no quantization support. - The loader of the ComfyUI node pack these files are made for (link to be added). The files hold the
native module's state dict, so they cannot be loaded with
transformersor the original Sapiens2 code as they are. - Input: a 1024 × 768 person crop (top-down, one person per box), normalized with ImageNet mean/std. The output is 308 keypoint heatmaps at 256 × 192, decoded with Sapiens2's UDP pipeline.
Speed and memory
These were measured on one RTX 3090 (24 GB) with PyTorch 2.11 + CUDA 12.8 (cu128), over a 233-frame 720 × 1280 clip, one person (one crop) per frame. The time covers crop, forward and decode of all 308 keypoints.
| File | s/frame | Peak VRAM | Load |
|---|---|---|---|
sapiens2_pose_1b_bf16 |
0.427 | 4.21 GiB | 4.4 s |
sapiens2_pose_1b_int8_convrot |
0.615 | 3.26 GiB | 3.3 s |
sapiens2_pose_5b_bf16 |
0.852 | 10.78 GiB | 7.1 s |
sapiens2_pose_5b_int8_convrot |
1.272 | 7.25 GiB | 4.1 s |
On this setup int8 is slower than bf16. comfy-kitchen's CUDA backend needs PyTorch built for CUDA 13.0 or newer (cu130+), so on cu128 the int8 layers run on its eager fallback. Here the gain from int8 is VRAM and disk, not speed. The 0.4b and 0.8b files were checked for parity but not timed on a full clip.
Licence
These weights are derived from Meta's Sapiens2 models
(facebookresearch/sapiens2, the facebook/sapiens2-pose-*
repos) and are distributed under the Sapiens2 License:
- Included: the licence is in this repository as
LICENSE.md, and it must be distributed with these files and any further derivative of them. - Use restrictions: by using or redistributing them you agree to its terms. Read the full text, especially the use restrictions in section 1.b.vi and the compliance requirements in sections 1.b.iii–v.
- Attribution: Sapiens2 and the original model weights are by Meta Platforms, Inc.; this repository only re-keys and quantizes them.
- Publications: under section 1.b.ii, research results obtained with these materials must acknowledge the use of Sapiens Materials in the publication.
- Warranty: the files are provided as is, without any warranty; see sections 3 and 4 of the licence.
- Downloads last month
- 17
Model tree for beycanai/sapiens2-convrot
Base model
facebook/sapiens2-pretrain-0.4b