Deep Reinforcement Learning and LLM Collaborative Architecture for QoS-Aware Traffic Engineering in 5G Core Networks

Authors

  • Sean D. Bryant Department of Computer Science, George Mason University, Fairfax, VA, USA. Author
  • Trevor A. Little School of Information Technology, University of Cincinnati, Cincinnati, OH, USA. Author
  • Benjamin Jarvinen Department of Computer Science, University of Alabama at Birmingham, Birmingham, AL, USA. Author
  • Vaibhav Ahuja Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS, USA. Author

Keywords:

Deep reinforcement learning, large language models, 5G core networks, traffic engineering, quality of service, collaborative intelligence, network slicing, governance, sustainability

Abstract

The proliferation of latency-sensitive, bandwidth-intensive, and ultra-reliable applications across fifth-generation mobile networks has fundamentally transformed the requirements placed upon traffic engineering in the core network. The 5G core, built on service-based architectures and network slicing, demands dynamic, fine-grained, and quality-of-service-aware routing decisions that operate at unprecedented scale and complexity. Traditional optimization heuristics and early machine learning approaches have struggled to maintain both performance and adaptability under such conditions. This paper presents a novel collaborative architecture that integrates deep reinforcement learning with large language models to achieve continuous, context-aware traffic engineering governed by explicit quality-of-service objectives. In this design, a deep reinforcement learning agent learns forwarding policies through interaction with a high-fidelity simulation of the core network, while a large language model functions as a semantic reasoning module, interpreting natural language operator intent, generating policy templates, and providing high-level guidance to constrain the reinforcement learning action space. The architecture is examined from system-level perspectives including modular composition, training stability, inference latency trade-offs, and the division of labor between learned control and symbolic reasoning. Particular attention is devoted to the implications of such a dual-model system for infrastructure sustainability, governance, fairness, and robustness in multi-tenant network environments. The paper further analyzes how collaborative intelligence architectures can absorb emerging traffic prediction models, referencing recent advances in spatial-temporal intelligence, and addresses deployment pathways that align with the ongoing softwarization of telecom infrastructure. Critical policy dimensions such as algorithmic transparency, slice isolation guarantees, and resource allocation equity are discussed as intrinsic design requirements rather than afterthoughts. The analysis suggests that hybrid architectures combining subsymbolic learning with the compositional reasoning capabilities of large language models represent a viable trajectory for the next generation of autonomous network management systems.

References

1. Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.

2. Xu, Z., Tang, J., Meng, J., Zhang, W., Wang, Y., Liu, C. H., & Yang, D. (2018). Experience-driven networking: A deep reinforcement learning based approach. IEEE Conference on Computer Communications, 1871–1879.

3. Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. ACM Workshop on Hot Topics in Networks, 50–56.

4. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

5. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

6. Manias, D. M., Shamim, I., Namuduri, S. R., & Shami, A. (2023). Large language models for autonomous network management: Opportunities and challenges. IEEE Communications Magazine, 61(10), 42–48.

7. Gonzalez, A. J., Ordonez-Lucena, J., Duque, D., & Lopez, D. (2024). Intent-based networking meets large language models: A new paradigm for network operations. IEEE Network, 38(2), 146–153.

8. Valadarsky, A., Schapira, M., Shahaf, D., & Tamar, A. (2017). Learning to route. ACM Workshop on Hot Topics in Networks, 185–191.

9. Li, R., Zhao, Z., Zhou, X., Ding, G., Chen, Y., Wang, Z., & Zhang, H. (2018). Intelligent 5G: When cellular networks meet artificial intelligence. IEEE Wireless Communications, 25(5), 10–17.

10. Chen, X., Li, Z., Zhang, Y., Long, R., Yu, H., & Leung, V. C. M. (2021). Reinforcement learning-based slice resource allocation for 5G core networks. IEEE Transactions on Network and Service Management, 18(4), 4562–4578.

11. Jiang, Y., Misra, V., & Rexford, J. (2024). Leveraging large language models for automated network troubleshooting. USENIX Symposium on Networked Systems Design and Implementation.

12. Gupta, S., Agrawal, A., Gopalakrishnan, K., & Narayanan, P. (2015). Deep learning with limited numerical precision. International Conference on Machine Learning, 1737–1746.

13. Bonati, L., Polese, M., D’Oro, S., Basagni, S., & Melodia, T. (2021). OpenRAN Gym: An open toolbox for data collection and experimentation with AI in O-RAN. IEEE International Conference on Computer Communications Workshops.

14. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2024). QLoRA: Efficient finetuning of quantized language models. Advances in Neural Information Processing Systems, 36.

15. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. International Conference on Learning Representations.

16. Konecny, J., McMahan, H. B., Yu, F. X., Richtarik, P., Suresh, A. T., & Bacon, D. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.

17. Preist, C., Shabajee, P., & Schien, D. (2016). Energy use in on-demand media delivery: The limits of current energy efficiency metrics. ACM SIGCAS Computers and Society, 46(3), 54–63.

18. Zhang, C., Zhang, H., Qiao, J., Li, Z., & Alouini, M. S. (2025). TIDES: Traffic Intelligence with DeepSeek Enhanced Spatial Temporal Prediction. IEEE Journal on Selected Areas in Communications.

19. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mane, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565.

20. Hogan, A., Blomqvist, E., Cochez, M., D’Amato, C., de Melo, G., Gutierrez, C., ... & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71.

21. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.

22. Feamster, N., & Rexford, J. (2018). Why the Internet needs verification: A view from the networking and security research communities. Communications of the ACM, 61(11), 56–63.

23. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.

24. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

25. Hong, C., Mandlekar, A., Xu, D., Tung, F., & Chi, E. (2023). Large language models are generalizable reasoning engines for task and motion planning. Conference on Robot Learning, 1714–1727.

Downloads

Published

2026-06-09

How to Cite

Deep Reinforcement Learning and LLM Collaborative Architecture for QoS-Aware Traffic Engineering in 5G Core Networks. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(1). https://ijaies.org/index.php/home/article/view/56