DriveAgent-WM: World Model-Driven Multimodal Autonomous Driving Agents with Chain-of-Thought Planning and Interactive Decision Making

Authors

  • Rmit Joshi Department of Computer Science, University of New Hampshire, Durham, NH, USA. Author
  • Hudson Graves Department of Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA. Author
  • Siddharth Bao Department of Computer Science, University of Central Florida, Orlando, FL, USA. Author

Keywords:

autonomous driving, world models, multimodal perception, chain-of-thought planning, interactive decision making, system architecture, AI governance

Abstract

The pursuit of safe and reliable autonomous driving has increasingly turned toward holistic architectures that integrate perception, planning, and control within a unified representational framework. This paper introduces DriveAgent-WM, a novel autonomous driving agent that leverages a world model-driven design to fuse multimodal sensor streams, perform explicit chain-of-thought planning, and enable interactive decision making under uncertain, dynamic traffic conditions. The system constructs an internal predictive model of the environment—a world model—that jointly encodes visual, LiDAR, radar, and trajectory data, enabling the agent to simulate future states and reason over possible action sequences. Chain-of-thought planning translates the latent predictions into interpretable reasoning steps, aligning the decision pipeline with human-like deliberation while preserving responsiveness. Interactive decision making is realized through closed-loop negotiation with other road users, governed by a policy optimization layer that continuously refines behavior based on simulated rollouts. The paper provides a comprehensive analysis of the architectural components, emphasizing the structural trade-offs between end-to-end learning and modular decomposition, the computational infrastructure required for real-time latent imagination, and the governance mechanisms necessary for safe deployment. Particular attention is directed toward robustness against distributional shift, fairness in interaction protocols, sustainability of compute-intensive world model training, and the broader policy implications of deploying cognitively autonomous agents on public roads. Through this systems-oriented lens, DriveAgent-WM is positioned not merely as a technical artifact but as a socio-technical construct that demands careful balancing of performance, interpretability, equity, and accountability.

References

1. Pomerleau, D. A. (1989). ALVINN: An autonomous land vehicle in a neural network. In D. S. Touretzky (Ed.), Advances in Neural Information Processing Systems 1 (pp. 305–313). Morgan Kaufmann.

2. Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., ... & Zieba, K. (2016). End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316.

3. Codevilla, F., Müller, M., López, A., Koltun, V., & Dosovitskiy, A. (2018). End-to-end driving via conditional imitation learning. In 2018 IEEE International Conference on Robotics and Automation (ICRA) (pp. 1–9). IEEE.

4. Ha, D., & Schmidhuber, J. (2018). World models. arXiv preprint arXiv:1803.10122.

5. Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., & Davidson, J. (2019). Learning latent dynamics for planning from pixels. In International Conference on Machine Learning (pp. 2555–2565). PMLR.

6. Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., ... & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11621–11631).

7. Liang, M., Yang, B., Chen, Y., Hu, R., & Urtasun, R. (2020). Multi-task multi-sensor fusion for 3D object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7345–7353).

8. Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2020). Dream to control: Learning behaviors by latent imagination. In International Conference on Learning Representations.

9. Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104.

10. van den Oord, A., Vinyals, O., & Kavukcuoglu, K. (2017). Neural discrete representation learning. In Advances in Neural Information Processing Systems (pp. 6306–6315).

11. Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., & Koltun, V. (2017). CARLA: An open urban driving simulator. In Conference on Robot Learning (pp. 1–16). PMLR.

12. He, K., Chen, X., Xie, S., Li, Y., Dollár, P., & Girshick, R. (2022). Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 16000–16009).

13. Jia, X., Yang, Z., Li, Q., Zhang, Z., & Yan, J. (2023). DriveGPT4: Interpretable end-to-end autonomous driving via large language model. arXiv preprint arXiv:2310.01412.

14. Chen, L., Wu, P., Chitta, K., Jaeger, B., Geiger, A., & Li, H. (2023). End-to-end autonomous driving: Challenges and frontiers. arXiv preprint arXiv:2306.16927.

15. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., ... & Le, Q. (2022). Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (pp. 24824–24837).

16. Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., ... & Zhou, D. (2023). Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations.

17. Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., ... & Li, H. (2023). Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 17853–17862).

18. Zhexiao Xiong, Xin Ye, Burhan Yaman, Sheng Cheng, Yiren Lu, Jingru Luo, Nathan Jacobs, and Liu Ren. UniDrive-WM: Unified understanding, planning and generation world model for autonomous driving. arXiv preprint arXiv:2601.04453, 2026.

19. Sun, L., Zhan, W., & Tomizuka, M. (2021). Interaction-aware planning for autonomous driving with game-theoretic reasoning. IEEE Transactions on Intelligent Vehicles, 6(2), 266–277.

20. Chai, Y., Sapp, B., Bansal, M., & Anguelov, D. (2019). MultiPath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conference on Robot Learning (pp. 86–99). PMLR.

21. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).

22. Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations.

23. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on Fairness, Accountability and Transparency (pp. 77–91). PMLR.

24. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1).

25. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.

Downloads

Published

2026-06-29

How to Cite

DriveAgent-WM: World Model-Driven Multimodal Autonomous Driving Agents with Chain-of-Thought Planning and Interactive Decision Making. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(1). https://ijaies.org/index.php/home/article/view/33