World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
Abstract
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
Community
🚀 We’re excited to release W²-VLA! 🎉
Task-conditioned future wrist modeling for fine-grained robot manipulation.
🌍 Global task context guides the prediction of task-relevant future wrist latents.
🧠 W²-CoT provides structured supervision to help shape the latent interface during training—without CoT decoding at inference.
📈 98.5% average success rate on LIBERO
🤖 60.71% Easy / 18.21% Hard on RoboTwin 2.0
⚡ Real-time action generation at over 80 Hz
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026)
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies (2026)
- DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation (2026)
- Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision (2026)
- HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation (2026)
- WAM4D: Fast 4D World Action Model via Spatial Register Tokens (2026)
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.05369 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper