PAI-Embedding Multilayer-Aligned-256 Pick Fruits Policy — 100-Epoch Target

This repository contains the audited final 100-epoch checkpoint for the three-way Pick Fruits downstream-policy comparison. Training is capped at this milestone; later 200-epoch states are not used by this study. The visual encoder is frozen; the Qwen3.5-0.8B policy and action head are trained. The trainable multilayer resampler is part of the policy checkpoint; it maps frozen PAI layers 7/14/21/28 to 3x256 policy tokens.

Artifact contract

  • Checkpoint: checkpoint/epoch_0100.pt
  • Step / epoch: 36,137 / 100
  • Encoder: PAI-Embedding
  • Encoder revision: 5d7e495e6e4b33b795d98b50588a06edef336a6f
  • Alignment: multilayer_aligned_256
  • Encoder layout: {"image_tokens_per_camera": 240, "num_cameras": 3, "representation": "cosmos3_layers_7_14_21_28_native_dense_240_per_camera+four_joint_fused_summaries;egodex_domain_3_zero_action_scaffold;trainable_resampler_to_3x256_plus_64;caip_compute_matched_849", "text_tokens": 4, "width": 1024}
  • Dataset: 100 Pick Fruits teleoperation episodes; 92,494 samples
  • Effective batch: 256 on 8 H100 GPUs
  • Comparison target: 100 epochs
  • Cosine-schedule horizon used by all three matched runs: 200 epochs; the common epoch-100 prefix is retained to avoid changing LR midway through aligned-256
  • Training loss at the epoch boundary: not available
  • Diagnostic: held-in diagnostic flow loss 0.02366 and MAE 0.00303 at step 36000
  • Schema: sha256:f60ed2905e168e731f7c6f9371b542bb0966a82392ab56ea8af28eca48525735
  • Normalization: sha256:bd52c50d4b22c8efafeb66054a82f4e9638a93cb4cd41009f406d5a140c9f2ad
  • Checkpoint SHA-256: 79c6df0358d4dae03b64bb331af4b9f66a5507676b7bbab643e9e0d95488bfb2
  • Training code: yunzeliu/VLA2Vec@7608781a451621fb40b7d3d89b5da07193bbaae5 on caip_downstream_policy

The diagnostic is computed on held-in training data. There is no held-out validation split, so it must not be reported as a generalization or robot success metric.

Export and inference

git clone --branch caip_downstream_policy https://github.com/yunzeliu/VLA2Vec.git
cd VLA2Vec
huggingface-cli download YunzeLiu/pai-embedding-aligned-256-pick-fruits-policy --local-dir /path/to/model
MODEL_DIR=/path/to/model

.venv/bin/python scripts/export_bundle.py   --checkpoint "$MODEL_DIR/checkpoint/epoch_0100.pt"   --out bundles/pai-embedding-aligned-256-pick-fruits-policy-epoch100   --encoder-pack assets/pai-embedding-encoder-pack --train-config "$MODEL_DIR/training/config.json"

HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1   .venv/bin/python scripts/serve.py   --bundle bundles/pai-embedding-aligned-256-pick-fruits-policy-epoch100   --encoder-pack assets/pai-embedding-encoder-pack --bind 'tcp://*:5678'

Validate the bundle with scripts/mock_client.py before hardware use. IK, collision checking, workspace limits and emergency-stop behavior remain the robot controller's responsibility.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train YunzeLiu/pai-embedding-aligned-256-pick-fruits-policy