Memory-Aware Remote Sensing Video Change Detection with Lightweight Vision Foundation Models
Keywords:
remote sensing, video change detection, vision foundation models, memory-aware systems, lightweight models, model deployment, AI governanceAbstract
Remote sensing video change detection is a critical task for environmental monitoring, urban planning, and disaster response, yet it imposes severe computational and memory demands due to high-resolution, temporally dense image sequences. The recent emergence of vision foundation models promises unprecedented segmentation and scene understanding capabilities, but their enormous parameter counts and reliance on attention mechanisms that scale quadratically with sequence length make direct deployment infeasible for resource-constrained satellite and airborne platforms. This paper investigates a systems-oriented framework that couples memory-aware sequence processing with lightweight adaptations of vision foundation models to enable sustainable, long-sequence video change detection. We discuss the architectural trade-offs between detection granularity, temporal coherence, memory footprint, and energy consumption, and we analyze how external memory banks, token compression, and hierarchical temporal aggregation can transform a heavy foundation model into a deployable edge-compatible system. The study further examines infrastructure-level considerations, including federated learning across distributed sensor swarms, on-satellite preprocessing, and carbon-aware scheduling. Governance dimensions such as fairness in land-cover change analysis, transparency of automated detection pipelines, and the geopolitical implications of ubiquitous overhead monitoring are foregrounded as integral to system design rather than afterthoughts. Throughout, we emphasize that the co-design of lightweight model architectures, memory management strategies, and policy-aware deployment workflows is essential for realizing the societal and scientific benefits of remote sensing video change detection at scale.
References
1. Shi, W., Zhang, M., Zhang, R., Chen, S., & Zhan, Z. (2020). Change detection based on artificial intelligence: State-of-the-art and challenges. ISPRS Journal of Photogrammetry and Remote Sensing, 169, 65–81.
2. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... & Girshick, R. (2023). Segment anything. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 4015–4026.
3. Daudt, R. C., Le Saux, B., & Boulch, A. (2018). Fully convolutional siamese networks for change detection. IEEE International Conference on Image Processing (ICIP), 4063–4067.
4. Chen, H., & Shi, Z. (2020). A spatial-temporal attention-based method and a new dataset for remote sensing image change detection. Remote Sensing, 12(10), 1662.
5. Oh, S. W., Lee, J. Y., Sunkavalli, K., & Kim, S. J. (2019). Video object segmentation using space-time memory networks. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9226–9235.
6. Mehta, S., & Rastegari, M. (2021). MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer. International Conference on Learning Representations (ICLR).
7. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. International Conference on Machine Learning (ICML), 6105–6114.
8. Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & Jégou, H. (2021). Training data-efficient image transformers & distillation through attention. International Conference on Machine Learning (ICML), 10347–10357.
9. Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., & Lillicrap, T. P. (2020). Compressive transformers for long-range sequence modelling. International Conference on Learning Representations (ICLR).
10. Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1290–1299.
11. Wang, X., Girshick, R., Gupta, A., & He, K. (2018). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7794–7803.
12. Ravi, N., Gabeur, V., Hu, Y. T., Hu, R., Ryali, C., Ma, T., ... & Feichtenhofer, C. (2024). SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714.
13. Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., & Luo, P. (2022). AdaptFormer: Adapting vision transformers for scalable visual recognition. Advances in Neural Information Processing Systems (NeurIPS), 35, 16664–16678.
14. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.
15. Howard, A., Sandler, M., Chu, G., Chen, L. C., Chen, B., Tan, M., ... & Adam, H. (2019). Searching for MobileNetV3. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1314–1324.
16. Tan, M., & Le, Q. V. (2021). EfficientNetV2: Smaller models and faster training. International Conference on Machine Learning (ICML), 10096–10106.
17. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations (ICLR).
18. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., ... & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2704–2713.
19. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L. M., Rothchild, D., ... & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
20. Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3), 50–60.
21. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1).
22. Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1(9), 389–399.
23. Kuffer, M., Pfeffer, K., & Sliuzas, R. (2016). Slums from space—15 years of slum mapping using remote sensing. Remote Sensing, 8(6), 455.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.