PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation
Abstract
World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.
Community
Project page: https://chrockey.github.io/PointWAM/
TL;DR: We forecast both the scene and the hands as 3D point trajectories, which lets us pre-train on human videos without any task-specific keypoint selection. Human-video pre-training alone gives +56.9pp on DexJoCo, and the model transfers to a real robot.
Happy to answer any questions!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation (2026)
- Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer (2026)
- GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments (2026)
- One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation (2026)
- GeomVLA: Unifying Scene, Motion, and Action in 3D (2026)
- Grounded Action Model: 3D Grounding as a Foundation for Robotics (2026)
- A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper