--- license: cc-by-4.0 arxiv: 2608.29434 datasets: - fafraob/point-cloud-reacher tags: - world-model - jepa - point-cloud - planning - robotics --- # Point-cloud world models: Reacher World-model checkpoints for the **Reacher** environment from the paper *Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals* ([arXiv:2608.29434](https://arxiv.org/abs/2608.29434)). Trained on [`fafraob/point-cloud-reacher`](https://huggingface.co/datasets/fafraob/point-cloud-reacher). Code, evaluation scripts and full details: [github.com/fafraob/point-lewm](https://github.com/fafraob/point-lewm). | folder | model | observation | |---|---|---| | `point-lewm/` | Point-LeWM (LeWM objective, PointViT encoder) | point cloud | | `point-delta-jepa/` | Point-Delta-JEPA (Delta-JEPA objective, PointViT encoder) | point cloud | | `image-delta-jepa/` | Delta-JEPA with a ViT-tiny image encoder trained from scratch | RGB | | `utonia-wm/` | JEPA predictor on frozen [Utonia](https://huggingface.co/Pointcept/Utonia) voxel features | point cloud | | `dino-wm/` | [DINO-WM](https://arxiv.org/abs/2411.04983) baseline, frozen DINOv2 ViT-S/14, no proprioception | RGB | Every folder except `dino-wm/` holds `weights.pt` (state dict) and `config.json` (Hydra instantiation spec) and loads with the `stable-worldmodel` folder loader from the code repository: ```python from huggingface_hub import snapshot_download import stable_worldmodel as swm local = snapshot_download("fafraob/point-cloud-reacher", allow_patterns=["point-lewm/*"]) # or any other folder model = swm.wm.utils.load_pretrained(f"{local}/point-lewm/") ``` `utonia-wm/` needs the public Utonia backbone (`utonia.pth` from `Pointcept/Utonia`, not redistributed here) placed at `checkpoints/utonia/utonia.pth` in the code repository. `dino-wm/` is in the layout the official [dino_wm](https://github.com/gaoyuezhou/dino_wm) code loads (`hydra.yaml`, `checkpoints/model_latest.pth`, `swm_h5_norm_stats.json` with the action normalisation); see the code repository for the evaluation wrapper. ## target2latent goal heads `point-lewm/target2latent/` and `point-delta-jepa/target2latent/` hold the typed-3-D-goal heads that map an object pose (cube position, T pose + ball position, agent position, or wrist + fingertip position) to the goal latent of the encoder they sit under, so planning needs no recorded goal observation. Four heads each: `mlp`, `mlp_z` (also conditioned on the current latent), `shortcut` (flow-matching shortcut model, samples in 1 to 64 steps), `shortcut_z`. Each folder holds one `model.pt` (state dict + architecture config); load with `target2latent.models.load_goal_model` from the code repository. ## Citation For the most recent citation see the [code repository](https://github.com/fafraob/point-lewm). ```bibtex @misc{oberweger2026doeslatentplanningsurvive, title = {Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals}, author = {Fabio F. Oberweger and Michael Schwingshackl and Markus Murschitz}, year = {2026}, eprint = {2608.29434}, archivePrefix = {arXiv}, primaryClass = {cs.LG}, url = {https://arxiv.org/abs/2608.29434}, } ```