Interactive World Models for Autonomous UAV Navigation and Real-Time Environmental Understanding in Complex Urban Scenarios
Keywords:
interactive world models; autonomous UAV navigation; urban environments; real-time perception; model-based reinforcement learning; edge computing; safety assuranceAbstract
Autonomous unmanned aerial vehicles operating in dense urban environments face a deeply coupled challenge: they must simultaneously construct rich spatiotemporal representations of their surroundings and act on those representations in real time under severe computational and safety constraints. Interactive world models, which learn to predict future sensory observations conditioned on potential actions, offer a promising system-level paradigm for bridging perception, planning, and control. This paper presents an interdisciplinary analysis of interactive world model architectures for UAV navigation, focusing on the structural trade-offs between model expressiveness, inference latency, and deployment robustness. We examine the integration of action-aware memory mechanisms that allow a UAV to simulate the outcomes of its own maneuvers within a learned latent dynamics space, thereby enabling anticipatory planning without exhaustive online search. The discussion extends to the supporting infrastructure required for edge-assisted model updates, the governance challenges of certifying learned world models, and the fairness implications of deploying such systems in diverse urban fabrics. Through a synthesis of advances in world models, model-based reinforcement learning, SLAM, semantic scene understanding, and safety assurance, the paper articulates a system-level research agenda that treats the world model not merely as a predictive component but as the central orchestrator of interactive autonomy. By foregrounding robustness, interpretability, and policy-aware design, we argue that interactive world models can serve as the integrative backbone for next-generation urban aerial systems.
References
1. Kendoul, F. (2012). Survey of advances in guidance, navigation, and control of unmanned rotorcraft systems. Journal of Field Robotics, 29(2), 315-378.
2. Ha, D., & Schmidhuber, J. (2018). Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems (pp. 2451-2462).
3. Thipphavong, D. P., Apaza, R., Barmore, B., Battiste, V., Burian, B., Dao, Q., ... & Swieringa, K. A. (2018). Urban air mobility airspace integration concepts and considerations. In 2018 Aviation Technology, Integration, and Operations Conference (p. 3676). AIAA.
4. Mur-Artal, R., & Tardós, J. D. (2017). ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras. IEEE Transactions on Robotics, 33(5), 1255-1262.
5. Geiger, A., Lenz, P., & Urtasun, R. (2012). Are we ready for autonomous driving? The KITTI vision benchmark suite. In 2012 IEEE Conference on Computer Vision and Pattern Recognition (pp. 3354-3361).
6. Xiong, Z., Song, Y., Kang, H., Yan, Q., Jiang, L., Yang, J., ... & Jacobs, N. (2026). ActWorld: From Explorable to Interactive World Model via Action-Aware Memory. arXiv preprint arXiv:2606.17730.
7. Chua, K., Calandra, R., McAllister, R., & Levine, S. (2018). Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Advances in Neural Information Processing Systems (pp. 4754-4765).
8. Yu, Z., Gong, Y., Gong, S., & Guo, Y. (2019). Edge intelligence for resource-constrained drone-based video surveillance. IEEE Network, 33(1), 128-135.
9. Fraichard, T., & Asama, H. (2004). Inevitable collision states — a step towards safer robots? Advanced Robotics, 18(10), 1001-1024.
10. Yuan, X., Wu, Q., Yu, D., & Zheng, J. (2019). A survey on deep reinforcement learning for UAV networks. IEEE Communications Surveys & Tutorials, 21(4), 3378-3415.
11. Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., & Levine, S. (2018). Visual foresight: Model-based deep reinforcement learning for vision-based robotic control. arXiv preprint arXiv:1812.00568.
12. Chen, Y. F., Liu, M., Everett, M., & How, J. P. (2017). Decentralized non-communicating multiagent collision avoidance with deep reinforcement learning. In 2017 IEEE International Conference on Robotics and Automation (ICRA) (pp. 285-292).
13. Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., ... & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11621-11631).
14. Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).
15. Quigley, M., Conley, K., Gerkey, B., Faust, J., Foote, T., Leibs, J., ... & Ng, A. Y. (2009). ROS: an open-source Robot Operating System. In ICRA workshop on open source software.
16. Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608.
17. Lim, W. Y. B., Luong, N. C., Hoang, D. T., Jiao, Y., Liang, Y. C., Yang, Q., ... & Miao, C. (2020). Federated learning in mobile edge networks: A comprehensive survey. IEEE Communications Surveys & Tutorials, 22(3), 2031-2063.
18. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565.
19. Crawford, K., & Calo, R. (2016). There is a blind spot in AI research. Nature, 538(7625), 311-313.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Data Intelligence and AI Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.