World ModelsarXiv preprint2025
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang et al.Tsinghua University, BAAI
摘要
We introduce DINO-WM, a world model built on top of pre-trained DINO visual features that enables zero-shot planning without task-specific training. By leveraging the rich semantic representations learned by DINO, our world model can predict future visual states and plan actions for novel tasks directly from visual observations. DINO-WM achieves strong zero-shot performance on manipulation and navigation tasks across diverse environments.
DINOworld modelpre-trained featureszero-shot planningvisual representation
技术细节
测试机器人
模型骨架
DINO-based World Model (Transformer)
编码器
DINOv2 ViT Encoder
解码器
Latent Dynamics Decoder + Planning Head
