Parameter-Efficient Foundation Model Adaptation for Medical Video Object Segmentation
Keywords:
foundation model adaptation, medical video segmentation, parameter-efficient fine-tuning, system architecture, governance, sustainability, fairnessAbstract
The proliferation of large foundation models has transformed video object segmentation, yet their direct application to medical imaging introduces a complex confluence of system-level constraints, data sensitivity, and domain-specific reliability requirements. While models such as the Segment Anything paradigm offer unprecedented zero-shot capacity, their computational scale impedes deployment in resource-limited clinical environments and raises concerns about long-term sustainability, fairness, and regulatory alignment. This paper presents a systems-oriented examination of parameter-efficient adaptation strategies tailored to medical video object segmentation. Rather than introducing a new algorithmic variation, we conduct a holistic analysis that spans infrastructure design, modular adaptation layers, memory footprint trade-offs, federated learning frameworks for privacy preservation, and governance mechanisms that ensure algorithmic fairness and explainability. The discussion integrates perspectives from large-scale distributed systems, edge computing, and health policy to illuminate structural interdependencies often overlooked in purely methodological contributions. We argue that sustainable deployment necessitates a multi-objective optimization that balances adaptation fidelity against computational cost, while embedding accountability through continuous monitoring and bias auditing. We further explore how parameter-efficient fine-tuning interacts with regulatory frameworks such as the European Union’s AI Act and the U.S. Food and Drug Administration’s evolving guidance on adaptive medical devices, linking technical design choices to compliance pathways. By synthesizing these dimensions, the paper provides a forward-looking roadmap for the academic community, clinicians, and system architects who must collaboratively steer the responsible integration of foundation models into real-world clinical video analysis workflows.
References
1. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... & Fei-Fei, L. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 4015-4026).
2. Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., ... & Misra, I. (2024). SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714.
3. Lialin, V., Muckatira, S., Shivagunde, N., & Rumshisky, A. (2023). Scaling down to scale up: A guide to parameter-efficient fine-tuning. arXiv preprint arXiv:2303.15647.
4. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., ... & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR).
5. Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., ... & Bengio, S. (2019). Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning (ICML) (pp. 2790-2799).
6. Ma, J., He, Y., Li, F., Han, L., You, C., & Wang, B. (2024). Segment anything in medical images. Nature Communications, 15(1), 654.
7. Cui, H., Mao, Y., Yang, Z., & Lo, B. (2023). SAM meets surgical video: Segment anything model for surgical instrument segmentation. arXiv preprint arXiv:2305.03579.
8. Oh, S. W., Lee, J. Y., Xu, N., & Kim, S. J. (2019). Video object segmentation using space-time memory networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 9226-9235).
9. Cheng, H. K., & Schwing, A. G. (2022). XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model. In European Conference on Computer Vision (ECCV) (pp. 640-658).
10. Liu, Y., Chen, K., Liu, C., Qin, Z., Luo, Z., & Wang, J. (2019). Structured knowledge distillation for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(3), 608-622.
11. Sheller, M. J., Edwards, B., Reina, G. A., Martin, J., Pati, S., Kotrotsou, A., ... & Bakas, S. (2020). Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Scientific Reports, 10(1), 12598.
12. Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4510-4520).
13. Bertasius, G., Wang, H., & Torresani, L. (2021). Is space-time attention all you need for video understanding? In International Conference on Learning Representations (ICLR).
14. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.
15. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453.
16. Vokinger, K. N., & Gasser, U. (2021). Regulating AI in medicine in the United States and Europe. The Lancet Digital Health, 3(6), e363-e364.
17. Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., ... & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82-115.
18. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645-3650).
19. Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations (ICLR).
20. Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., & Wermter, S. (2019). Continual lifelong learning with neural networks: A review. Neural Networks, 113, 54-71.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.