Jiang, Fan, Sun, Zhaoxu, Wang, Mengchao, Zhu, Ziyu, Wang, Chiyu, Zhang, Yunpeng, Liu, Wenlin, Wang, Yun, Zheng, Xue, Sun, Rui, Ni, Junfeng, Pan, Hongyu, Sun, Zhongxu, Yu, Fei, Ge, Zengye, Du, Mengmeng, Fan, Nianfei, Sun, Mingchao, Liu, Yu, Yongchang, Zhu, Yanqing, Wang, Jiahang, Ying, Ning, Xuan, Yuze, Yang, Di, Liu, Zhicheng, Gao, Zhe, Xu, Tingbing, Sui, Jiacheng, Yang, Wenjin, Lai, Junnan, Liu, Shufeng, Liu, Yuan, Zhou, Zheng, Peng, Yingliang, Cao, Dawei, Sheng, Kaifeng, Cai, Yuxiang, Lu, Fei, Xu, Mu, Guo, Ning
Abstract
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
Chinese Translation
我们提出了ABot-World-0,这是一种基于动作条件的视频世界模型,旨在实现实时、长时间闭环交互,支持的多源数据基础设施涵盖AAA游戏、仿真引擎和互联网视频,以学习可控的世界动态。WorldExplorer通过训练反馈指导的代理驱动收集,而统一管道应用14项确定性质量检查、基于VLM的评估以及同步的动作和文本注释。我们通过教师强制和常微分方程(ODE)蒸馏,逐步将双向动作条件教师蒸馏为因果学生,并引入LongForcing以将长时间学生自展开与扩展视野教师对齐,从而减轻累积的分布偏移和自回归漂移。原始键盘动作提供了一个统一的控制接口,用于场景漫游和第三人称角色交互,而参考角色记忆则在第三人称展开过程中提供持久的外观线索,以确保身份一致性。在部署方面,我们共同设计了一个流式推理堆栈,配备轻量级的变分自编码器(VAE)解码器、高效的注意力机制、内存感知调度和低比特的DiT推理。在优化的低比特配置下,ABot-World-0在单个NVIDIA RTX 5090桌面GPU上以高达16帧每秒的速度流式传输720P视频,具有1.2秒的动作到首帧延迟和约19GiB的峰值显存。对WorldRoamBench和扩展交互展开的实验表明了竞争性的可控性和一致的长时间世界演变。