Transformer-Based Spatio-Temporal Attention Optimization for Real-Time Autonomous Driving Video Segmentation
Keywords:
autonomous driving, video segmentation, vision transformer, spatio-temporal attention, real-time systems, edge computing, algorithmic fairnessAbstract
Real-time video segmentation constitutes a fundamental perception capability for autonomous driving systems, enabling fine-grained scene understanding under stringent latency constraints. The emergence of vision transformer architectures has introduced powerful global context modeling, yet their computational intensity poses significant challenges for deployment on resource-constrained vehicular platforms. This paper presents a comprehensive systems-level analysis of transformer-based spatio-temporal attention optimization for autonomous driving video segmentation, examining the structural trade-offs, architectural design principles, and infrastructural requirements that govern the transition from laboratory benchmarks to safety-critical real-world operation. We investigate the interplay between spatial attention precision, temporal coherence maintenance, and inference latency within embedded computing environments, while addressing broader considerations of robustness, fairness, governance, and sustainability. The discussion encompasses hardware-aware model design, distributed edge-cloud processing topologies, adversarial resilience mechanisms, and algorithmic equity across diverse operational design domains. Through a synthesis of contemporary research and engineering practice, we identify critical pathways for achieving the reliability, efficiency, and ethical alignment necessary for large-scale autonomous deployment. The paper further explores policy frameworks, standardization imperatives, and the societal dimensions of perceptual artificial intelligence in mobility systems, offering a multidisciplinary perspective that bridges systems engineering with socio-technical governance.
References
1. Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). SegNet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481–2495.
2. Oh, S. W., Lee, J.-Y., Xu, N., & Kim, S. J. (2019). Video object segmentation using space-time memory networks. Proceedings of the IEEE International Conference on Computer Vision, 9226–9235.
3. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
4. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE International Conference on Computer Vision, 10012–10022.
5. Bertasius, G., Wang, H., & Torresani, L. (2021). Is space-time attention all you need for video understanding? Proceedings of the International Conference on Machine Learning.
6. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.
7. Chen, W., Huang, W., Du, X., Song, X., Wang, Z., & Zhou, D. (2023). Hardware-aware efficient vision transformer design for autonomous driving. IEEE Transactions on Circuits and Systems for Video Technology, 33(8), 3845–3858.
8. Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., & Blankevoort, T. (2021). A white paper on neural network quantization. arXiv preprint arXiv:2106.08295.
9. Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39.
10. Koopman, P., & Wagner, M. (2017). Autonomous vehicle safety: An interdisciplinary challenge. IEEE Intelligent Transportation Systems Magazine, 9(1), 90–96.
11. Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. International Conference on Learning Representations.
12. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of the Conference on Fairness, Accountability and Transparency, 77–91.
13. Daily, M., Medasani, S., Behringer, R., & Trivedi, M. (2017). Self-driving cars. Computer, 50(12), 18–23.
14. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
15. Kalra, N., & Paddock, S. M. (2016). Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability? Transportation Research Part A: Policy and Practice, 94, 182–193.
16. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.
17. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778.
18. Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1290–1299.
19. Xu, N., Yang, L., Fan, Y., Yue, D., Liang, Y., Yang, J., & Huang, T. (2018). YouTube-VOS: A large-scale video object segmentation benchmark. arXiv preprint arXiv:1809.03327.
20. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. Proceedings of the IEEE International Conference on Computer Vision.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.