Deep Reinforcement Learning and LLM Collaborative Architecture for QoS-Aware Traffic Engineering in 5G Core Networks
Keywords:
Deep reinforcement learning, large language models, 5G core networks, traffic engineering, quality of service, collaborative intelligence, network slicing, governance, sustainabilityAbstract
The proliferation of latency-sensitive, bandwidth-intensive, and ultra-reliable applications across fifth-generation mobile networks has fundamentally transformed the requirements placed upon traffic engineering in the core network. The 5G core, built on service-based architectures and network slicing, demands dynamic, fine-grained, and quality-of-service-aware routing decisions that operate at unprecedented scale and complexity. Traditional optimization heuristics and early machine learning approaches have struggled to maintain both performance and adaptability under such conditions. This paper presents a novel collaborative architecture that integrates deep reinforcement learning with large language models to achieve continuous, context-aware traffic engineering governed by explicit quality-of-service objectives. In this design, a deep reinforcement learning agent learns forwarding policies through interaction with a high-fidelity simulation of the core network, while a large language model functions as a semantic reasoning module, interpreting natural language operator intent, generating policy templates, and providing high-level guidance to constrain the reinforcement learning action space. The architecture is examined from system-level perspectives including modular composition, training stability, inference latency trade-offs, and the division of labor between learned control and symbolic reasoning. Particular attention is devoted to the implications of such a dual-model system for infrastructure sustainability, governance, fairness, and robustness in multi-tenant network environments. The paper further analyzes how collaborative intelligence architectures can absorb emerging traffic prediction models, referencing recent advances in spatial-temporal intelligence, and addresses deployment pathways that align with the ongoing softwarization of telecom infrastructure. Critical policy dimensions such as algorithmic transparency, slice isolation guarantees, and resource allocation equity are discussed as intrinsic design requirements rather than afterthoughts. The analysis suggests that hybrid architectures combining subsymbolic learning with the compositional reasoning capabilities of large language models represent a viable trajectory for the next generation of autonomous network management systems.
References
1. Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.
2. Xu, Z., Tang, J., Meng, J., Zhang, W., Wang, Y., Liu, C. H., & Yang, D. (2018). Experience-driven networking: A deep reinforcement learning based approach. IEEE Conference on Computer Communications, 1871–1879.
3. Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. ACM Workshop on Hot Topics in Networks, 50–56.
4. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
5. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.
6. Manias, D. M., Shamim, I., Namuduri, S. R., & Shami, A. (2023). Large language models for autonomous network management: Opportunities and challenges. IEEE Communications Magazine, 61(10), 42–48.
7. Gonzalez, A. J., Ordonez-Lucena, J., Duque, D., & Lopez, D. (2024). Intent-based networking meets large language models: A new paradigm for network operations. IEEE Network, 38(2), 146–153.
8. Valadarsky, A., Schapira, M., Shahaf, D., & Tamar, A. (2017). Learning to route. ACM Workshop on Hot Topics in Networks, 185–191.
9. Li, R., Zhao, Z., Zhou, X., Ding, G., Chen, Y., Wang, Z., & Zhang, H. (2018). Intelligent 5G: When cellular networks meet artificial intelligence. IEEE Wireless Communications, 25(5), 10–17.
10. Chen, X., Li, Z., Zhang, Y., Long, R., Yu, H., & Leung, V. C. M. (2021). Reinforcement learning-based slice resource allocation for 5G core networks. IEEE Transactions on Network and Service Management, 18(4), 4562–4578.
11. Jiang, Y., Misra, V., & Rexford, J. (2024). Leveraging large language models for automated network troubleshooting. USENIX Symposium on Networked Systems Design and Implementation.
12. Gupta, S., Agrawal, A., Gopalakrishnan, K., & Narayanan, P. (2015). Deep learning with limited numerical precision. International Conference on Machine Learning, 1737–1746.
13. Bonati, L., Polese, M., D’Oro, S., Basagni, S., & Melodia, T. (2021). OpenRAN Gym: An open toolbox for data collection and experimentation with AI in O-RAN. IEEE International Conference on Computer Communications Workshops.
14. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2024). QLoRA: Efficient finetuning of quantized language models. Advances in Neural Information Processing Systems, 36.
15. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. International Conference on Learning Representations.
16. Konecny, J., McMahan, H. B., Yu, F. X., Richtarik, P., Suresh, A. T., & Bacon, D. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.
17. Preist, C., Shabajee, P., & Schien, D. (2016). Energy use in on-demand media delivery: The limits of current energy efficiency metrics. ACM SIGCAS Computers and Society, 46(3), 54–63.
18. Zhang, C., Zhang, H., Qiao, J., Li, Z., & Alouini, M. S. (2025). TIDES: Traffic Intelligence with DeepSeek Enhanced Spatial Temporal Prediction. IEEE Journal on Selected Areas in Communications.
19. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mane, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565.
20. Hogan, A., Blomqvist, E., Cochez, M., D’Amato, C., de Melo, G., Gutierrez, C., ... & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71.
21. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.
22. Feamster, N., & Rexford, J. (2018). Why the Internet needs verification: A view from the networking and security research communities. Communications of the ACM, 61(11), 56–63.
23. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
24. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
25. Hong, C., Mandlekar, A., Xu, D., Tung, F., & Chi, E. (2023). Large language models are generalizable reasoning engines for task and motion planning. Conference on Robot Learning, 1714–1727.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.