Learning Physical Interaction Dynamics from Visual Experience for General-Purpose Robotic Manipulation

Authors

  • Jase Haffmen School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, OR, USA. Author
  • Metthias Bndrews Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO, USA. Author
  • Raj D. Pethaeak School of Information Technology, University of Cincinnati, Cincinnati, OH, USA. Author

Keywords:

physical interaction dynamics; visual robot learning; predictive world models; robotic manipulation; simulation-to-real transfer; robustness; governance

Abstract

General-purpose robotic manipulation requires reliable prediction of how objects move, deform, slip, and interact under applied forces. Visual experience is an especially scalable source of supervision because cameras are inexpensive and observations can be collected continuously during robot operation. However, converting raw visual streams into actionable physical interaction dynamics remains a fundamental systems problem. This article examines the design of robot learning systems that acquire manipulation dynamics from visual experience. We address the structural trade-offs between end-to-end visuomotor control and model-based predictive world models, with particular attention to action-conditioned prediction, memory, latent state representation, and temporal consistency. The discussion spans simulation infrastructure, real-world data collection, cross-embodiment transfer, robustness under distribution shift, safety, fairness, sustainability, and governance. We argue that accurate physical prediction alone is insufficient for general-purpose deployment. The surrounding data infrastructure, evaluation protocols, regulatory mechanisms, and organizational practices are equally important. This article provides a critical synthesis of current approaches and outlines directions for building more accountable, robust, and scalable visually grounded manipulation systems.

References

1. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436-444.

2. Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. Journal of Machine Learning Research, 17(1), 1334-1373.

3. Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the 34th International Conference on Machine Learning, 1126-1135.

4. Ha, D., & Schmidhuber, J. (2018). World models. arXiv preprint arXiv:1803.10122.

5. Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., & Davidson, J. (2019). Learning latent dynamics for planning from pixels. Proceedings of the 36th International Conference on Machine Learning, 2555-2565.

6. Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., ... & Silver, D. (2020). Mastering Atari, Go, chess and shogi by planning with a learned model. Nature, 588(7839), 604-609.

7. Ebert, F., Finn, C., Lee, A. X., & Levine, S. (2017). Self-supervised visual planning with temporal skip connections. Proceedings of the 1st Annual Conference on Robot Learning, 344-356.

8. Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., ... & Hassabis, D. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354-359.

9. Xiong, Zhexiao, et al. "ActWorld: From Explorable to Interactive World Model via Action-Aware Memory." arXiv preprint arXiv:2606.17730 (2026).

10. Coumans, E., & Bai, Y. (2016). PyBullet, a Python module for physics simulation for games, robotics and machine learning. http://pybullet.org

11. Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A physics engine for model-based control. Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 5026-5033.

12. Zeng, A., Song, S., Lee, J., Rodriguez, A., & Funkhouser, T. (2020). TossingBot: Learning to throw arbitrary objects with residual physics. IEEE Transactions on Robotics, 36(4), 1307-1319.

13. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., ... & Zitkovich, B. (2023). RT-1: Robotics Transformer for real-world control at scale. arXiv preprint arXiv:2212.06817.

14. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., ... & Zitkovich, B. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818.

15. Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., ... & Shah, R. (2023). Open X-Embodiment: Robotic learning datasets and RT-X models. arXiv preprint arXiv:2310.08864.

16. Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., ... & de Freitas, N. (2022). A generalist agent. arXiv preprint arXiv:2205.06175.

17. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning, 8748-8763.

18. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

19. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., ... & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220-229.

20. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of the Conference on Fairness, Accountability, and Transparency, 77-91.

21. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565.

22. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R. B., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

23. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1). https://doi.org/10.1162/99608f92.8cd550d1

24. Crawford, K., & Paglen, T. (2021). Excavating AI: The politics of images in machine learning training sets. AI & Society, 36(4), 1105-1116.

Downloads

Published

2026-07-29

How to Cite

Learning Physical Interaction Dynamics from Visual Experience for General-Purpose Robotic Manipulation. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(2). https://ijaies.org/index.php/home/article/view/106