Memory-Enhanced Human Action Parsing and Object Interaction Segmentation in Complex Video Streams

Authors

  • Jingwendong Zhu School of Computing, Clemson University, Clemson, SC, USA. Author
  • Anil Bhakla Department of Computer Science, University of Alabama at Birmingham, Birmingham, AL, USA. Author

Keywords:

video understanding, memory-augmented networks, human-object interaction, action segmentation, system architecture, fairness, edge computing

Abstract

The proliferation of high-resolution surveillance, egocentric, and broadcast video data has exposed fundamental limitations in conventional architectures for human action parsing and object interaction segmentation. Short temporal windows, isolated treatment of action and object cues, and static model designs fail to capture the long-range dependencies, occlusions, and compositional dynamics that characterize complex real-world streams. This paper presents a comprehensive systems-oriented analysis of memory-enhanced frameworks that jointly model human actions and object interactions across extended temporal horizons. We argue that the integration of explicit memory modules, which can store, update, and selectively retrieve spatiotemporal context, constitutes a structural transformation rather than a mere accuracy improvement. The discussion dissects architectural trade-offs among fixed-size latent memory pools, self-attentive memory banks, and dynamic external memory controllers, linking design choices to compute cost, latency, and the capacity to sustain coherent object-action associations over thousands of frames. By embedding the technical analysis within a larger socio-technical perspective, the paper examines infrastructure requirements, edge-cloud partitioning strategies, energy sustainability, data governance, and the fairness implications of action recognition systems deployed in diverse human environments. Policy recommendations for lifecycle monitoring, bias auditing, and privacy-preserving inference are distilled from case illustrations in public safety and assistive technology domains. The paper concludes by identifying open systems challenges, including memory staleness, cross-modal alignment under distribution shift, and the need for standardized benchmarks that evaluate consistency rather than frame-level accuracy alone. The synthesis positions memory-enhanced video parsing as a critical enabler of trustworthy, scalable, and context-aware perception systems.

References

1. Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6299-6308). https://doi.org/10.1109/CVPR.2017.502

2. Feichtenhofer, C., Fan, H., Malik, J., & He, K. (2019). SlowFast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 6202-6211).

3. Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., ... & Hassabis, D. (2016). Hybrid computing using a neural network with dynamic external memory. Nature, 538(7626), 471-476.

4. Oh, S. W., Lee, J.-Y., Sunkavalli, K., & Kim, C.-S. (2019). Fast user-guided video object segmentation by interaction-and-propagation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5247-5256).

5. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.

6. Gkioxari, G., Girshick, R., Dollár, P., & He, K. (2018). Detecting and recognizing human-object interactions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 8359-8367).

7. Cheng, H.-K., & Schwing, A. G. (2022). XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model. In European Conference on Computer Vision (pp. 640-658). Springer.

8. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645-3650).

9. Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., ... & Schafer, B. (2018). AI4People—an ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and Machines, 28(4), 689-707.

10. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (pp. 77-91). PMLR.

11. Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. In International Conference on Learning Representations.

12. Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van Gool, L. (2016). Temporal segment networks: Towards good practices for deep action recognition. In European Conference on Computer Vision (pp. 20-36). Springer.

13. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015-4026).

14. Mao, H., Cheng, J., Yang, J., Jayagopi, D., Gu, L., Ye, A., ... & Fei-Fei, L. (2023). Ego4D: A large-scale egocentric video dataset and benchmark suite. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(6), 7039-7056.

15. Zhang, H., Ananthanarayanan, G., Bahl, P., Philipose, M., & Gibbons, P. B. (2017). Live video analytics at scale with approximation and delay-tolerance. In 14th USENIX Symposium on Networked Systems Design and Implementation (pp. 377-392).

16. Padmanabhan, V. N., Ramjee, R., Srikanth, S., & Seetharam, A. (2018). The case for edge-based video analytics. In Proceedings of the ACM Workshop on Internet of Things and Edge Computing (pp. 1-6).

17. Kosmopoulos, D., Petropoulos, T., & Argyros, A. (2022). A survey on human action recognition in videos: Challenges, datasets, and methods. ACM Computing Surveys, 55(2), 1-47.

18. Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., & Sorkine-Hornung, A. (2016). A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 724-732).

19. Ravi, N., Gabeur, V., Hu, J., Verbeek, J., & Zisserman, A. (2021). Video action classification with text supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 12905-12915).

20. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.

Downloads

Published

2026-06-13

How to Cite

Memory-Enhanced Human Action Parsing and Object Interaction Segmentation in Complex Video Streams. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(1). https://ijaies.org/index.php/home/article/view/26