AI-Driven Resource Scheduling and Optimization in Distributed Computing Environments
Keywords:
AI-driven scheduling, distributed computing, resource optimization, deep reinforcement learning, cloud computing, edge computing, fairness, sustainabilityAbstract
The proliferation of cloud computing, Internet of Things ecosystems, and large-scale machine learning workloads has placed unprecedented demands on distributed computing infrastructures, compelling a fundamental rethinking of resource scheduling and optimization. Traditional heuristic and rule-based schedulers, while performant in static or moderately dynamic settings, prove inadequate when confronted with heterogeneous hardware, multi-tenant volatility, and stringent multi-objective requirements spanning latency, throughput, energy efficiency, and fairness. This paper examines the emergence of AI-driven scheduling paradigms—encompassing deep reinforcement learning, meta-learning, and online learning—that dynamically orchestrate computational, storage, and network resources across geo-distributed clusters, edge nodes, and federated data centers. The paper provides a system-level analysis, addressing architectural foundations of distributed resource management frameworks, the structural trade-offs introduced by AI integration, and the operational realities of deployment at scale. Rather than focusing on algorithmic minutiae, the discussion foregrounds the interplay between autonomy and over-provisioning, the governance challenges arising from opaque decision processes, and the socio-technical dimensions of fairness, sustainability, and regulatory accountability. Through a cross-domain synthesis of cloud orchestration, edge intelligence, and carbon-aware computing, the paper delineates how AI-driven schedulers can simultaneously pursue utilization efficiency, service-level agreement guarantees, and environmental responsibility. Critical reflections on robustness, adversarial vulnerability, and policy alignment underscore that responsible AI scheduling demands not only technical sophistication but also institutional mechanisms for auditability, recourse, and equitable resource distribution. The analysis culminates in a forward-looking agenda that couples adaptive infrastructure with transparent governance, advocating for systems that learn continuously while remaining accountable to human values.
References
1. Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., ... & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58.
2. Jennings, B., & Stadler, R. (2015). Resource management in clouds: Survey and research challenges. Journal of Network and Systems Management, 23(3), 567–619.
3. Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., & Wilkes, J. (2015). Large-scale cluster management at Google with Borg. Proceedings of the Tenth European Conference on Computer Systems (EuroSys), 1–17.
4. Hindman, B., Konwinski, A., Zaharia, M., Ghodsi, A., Joseph, A. D., Katz, R., ... & Stoica, I. (2011). Mesos: A platform for fine-grained resource sharing in the data center. Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI), 295–308.
5. Vavilapalli, V. K., Murthy, A. C., Douglas, C., Agarwal, S., Konar, M., Evans, R., ... & Baldeschwieler, E. (2013). Apache Hadoop YARN: Yet another resource negotiator. Proceedings of the 4th Annual Symposium on Cloud Computing (SOCC), 1–16.
6. Burns, B., Grant, B., Oppenheimer, D., Brewer, E., & Wilkes, J. (2016). Borg, Omega, and Kubernetes. Communications of the ACM, 59(5), 50–57.
7. Schwarzkopf, M., Konwinski, A., Abd-El-Malek, M., & Wilkes, J. (2013). Omega: Flexible, scalable schedulers for large compute clusters. Proceedings of the 8th ACM European Conference on Computer Systems (EuroSys), 351–364.
8. Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
9. Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. Proceedings of the 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud), 10–10.
10. Isard, M., Budiu, M., Yu, Y., Birrell, A., & Fetterly, D. (2007). Dryad: Distributed data-parallel programs from sequential building blocks. Proceedings of the 2nd ACM SIGOPS/EuroSys European Conference on Computer Systems, 59–72.
11. Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., ... & Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 265–283.
12. Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., ... & Stoica, I. (2018). Ray: A distributed framework for emerging AI applications. Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 561–577.
13. Delimitrou, C., & Kozyrakis, C. (2013). Paragon: QoS-aware scheduling for heterogeneous datacenters. Proceedings of the 18th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 77–88.
14. Delimitrou, C., & Kozyrakis, C. (2014). Quasar: Resource-efficient and QoS-aware cluster management. Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 127–144.
15. Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518, 529–533.
16. Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. Proceedings of the 15th ACM Workshop on Hot Topics in Networks (HotNets), 50–56.
17. Peng, Y., Bao, Y., Chen, Y., Wu, C., & Guo, C. (2018). Optimus: An efficient dynamic resource scheduler for deep learning clusters. Proceedings of the Thirteenth European Conference on Computer Systems (EuroSys), 1–14.
18. Ghodsi, A., Zaharia, M., Hindman, B., Konwinski, A., Shenker, S., & Stoica, I. (2011). Dominant resource fairness: Fair allocation of multiple resource types. Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI), 323–336.
19. Masanet, E., Shehabi, A., Lei, N., Smith, S., & Koomey, J. (2020). Recalibrating global data center energy-use estimates. Science, 367(6481), 984–986.
20. Radovanovic, A., Koningstein, R., Schneider, I., Chen, B., Duarte, A., Roy, B., ... & Cirne, W. (2022). Carbon-aware computing for datacenters. IEEE Micro, 42(4), 14–25.
21. Wiesner, P., Beier, I., Esterle, L., Müller, H., & Schill, A. (2021). Let’s wait awhile: How temporal workload shifting can reduce carbon emissions in the cloud. Proceedings of the 22nd International Middleware Conference (Middleware), 260–272.
22. Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., ... & Schafer, B. (2018). AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and Machines, 28(4), 689–707.
23. Kroll, J. A., Huey, J., Barocas, S., Felten, E. W., Reidenberg, J. R., Robinson, D. G., & Yu, H. (2017). Accountable algorithms. University of Pennsylvania Law Review, 165, 633–705.
24. Shahrad, M., Fonseca, R., Goiri, I., Chaudhry, G., Batum, P., Cooke, J., ... & Bianchini, R. (2020). Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. Proceedings of the 2020 USENIX Annual Technical Conference (ATC), 205–218.
25. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT), 59–68.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.