Memory-Guided Video Scene Understanding for Embodied AI Navigation and Interactive Perception
Keywords:
embodied AI, video scene understanding, memory architectures, visual navigation, interactive perception, system design, fairness, governanceAbstract
Embodied artificial intelligence systems that operate in dynamic real-world environments must continuously parse streaming visual information, retain context over extended temporal horizons, and adapt their behavior in response to changing scenes. This paper presents a system-level analysis of memory-guided video scene understanding as a foundational capability for embodied navigation and interactive perception. We examine the architectural space of memory mechanisms—ranging from differentiable external stores and episodic controllers to transformer-based long-range context models—and trace their integration into video understanding pipelines that serve navigation, manipulation, and multimodal reasoning. Rather than focusing on algorithmic novelty, the discussion foregrounds structural trade-offs in capacity, retrieval latency, training stability, and deployment footprint, and it situates these technical considerations within broader sociotechnical frameworks. The paper dissects the interplay between spatial-topological mapping, object-centric memory, and continual adaptation in navigation tasks, and it extends the analysis to interactive perception where memory of past interactions shapes affordance learning and task planning. System-level challenges are then explored, including edge-cloud partitioning, energy efficiency, robustness under sensor degradation, and the tension between persistent memory and privacy. Governance, fairness, and safety implications are addressed through the lenses of bias auditing, regulatory compliance, and inclusive design, emphasizing that memory-guided embodied agents straddle the boundary between high-stakes autonomy and socially embedded operation. By synthesizing architectural insights with deployment and policy perspectives, the paper provides a holistic reference for researchers, system architects, and policymakers engaged in the development of trustworthy embodied AI.
References
1. Weston, J., Chopra, S., & Bordes, A. (2015). Memory networks. In Proceedings of the International Conference on Learning Representations (ICLR).
2. Graves, A., Wayne, G., & Danihelka, I. (2014). Neural Turing machines. arXiv preprint arXiv:1410.5401.
3. Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., & Salakhutdinov, R. (2019). Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
4. Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., & Blundell, C. (2017). Neural episodic control. In Proceedings of the 34th International Conference on Machine Learning (ICML).
5. Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., & Batra, D. (2019). Habitat: A platform for embodied AI research. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
6. Xia, F., Zamir, A. R., He, Z., Sax, A., Malik, J., & Savarese, S. (2018). Gibson Env: Real-world perception for embodied agents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
7. Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., & Zhang, Y. (2017). Matterport3D: Learning from RGB-D data in indoor environments. In Proceedings of International Conference on 3D Vision (3DV).
8. Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., & van den Hengel, A. (2018). Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
9. Chaplot, D. S., Gandhi, D., Gupta, S., Gupta, A., & Salakhutdinov, R. (2020). Learning to explore using active neural SLAM. In Proceedings of the International Conference on Learning Representations (ICLR).
10. Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Gordon, D., Zhu, Y., Gupta, A., & Farhadi, A. (2017). AI2-THOR: An interactive 3D environment for visual AI. arXiv preprint arXiv:1712.05474.
11. Gan, C., Schwartz, J., Alter, S., Schrimpf, M., Traer, J., De Freitas, J., Kubilius, J., Bhandwaldar, A., Haber, N., Sano, M., Kim, K., Wang, E., Mrowca, D., Lingelbach, M., Curtis, A., Feigelis, K., Bear, D. M., Gutfreund, D., Cox, D. D., … Yamins, D. L. K. (2020). The threeDWorld transport challenge: A physically realistic simulation for embodied AI. In Advances in Neural Information Processing Systems (NeurIPS).
12. Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., Savva, M., Zhao, Y., & Batra, D. (2021). Habitat 2.0: Training home assistants to rearrange their habitat. In Advances in Neural Information Processing Systems (NeurIPS).
13. Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., & Batra, D. (2020). DD-PPO: Learning near-perfect pointgoal navigators from 2.5 billion frames. In Proceedings of the International Conference on Learning Representations (ICLR).
14. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
15. Yang, J., Gao, M., Li, Z., Gao, S., Wang, F., & Zheng, F. (2024). SAM 2: The next generation of segment anything for image and video. arXiv preprint arXiv:2408.00714.
16. Minderer, M., Gritsenko, A., Stone, A., Neumann, M., Weissenborn, D., Dosovitskiy, A., Mahendran, A., Arnab, A., Dehghani, M., Shen, Z., Wang, X., Zhai, X., Kipf, T., & Houlsby, N. (2022). Simple open-vocabulary object detection with vision transformers. In Proceedings of the European Conference on Computer Vision (ECCV).
17. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
18. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.
19. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT).
20. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT).
21. Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT).
22. Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., & Srinivas, A. (2020). Reinforcement learning with augmented data. In Advances in Neural Information Processing Systems (NeurIPS).
23. Howard, A., & Borenstein, J. (2018). The ugly truth about ourselves and our robot creations: The problem of bias and social inequity. Science and Engineering Ethics, 24(5), 1521–1536.
24. European Commission. (2021). Proposal for a Regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). COM(2021) 206 final.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.