An Intelligent Predictive GPU Scheduling Framework for Deep Learning Workloads in Large-Scale Cloud Environments
DOI:
https://doi.org/10.15680/IJCTECE.2025.0806025Keywords:
GPU Scheduling, Deep Learning, Cloud Computing, Resource Management, Predictive Framework, Machine Learning, Optimization AlgorithmsAbstract
The paper is about the problems of optimally scheduling the resources of GPUs when it comes to large-scale deep learning workloads within the cloud infrastructure. Due to the fact that the field of deep learning requires a noticeably large amount of computational power, GPUs play a crucial role in hastening these operations. Nevertheless, the management of the GPU resources in the cloud is still complicated because of the dynamics of the workload and resources allocation that requires efficient implementation. This paper introduces a predictive graphic card scheduling system, which takes advantage of machine learning to predict resource demands through the nature of workload. The architecture combines workload prediction, the control of the GPU resources, and optimization algorithms to allocate resources in advance that the deep learning tasks get the required amount of GIS in time and at the same time, the active time is minimized as well as the contention of the resources
The framework relies on the historical performance data as well as analysis of workload behavior to forecast future demands which are then adjusted in terms of scheduling strategies. The strategy not only maximizes the use of the GPU but also leads to the overall performance of cloud-based systems since it minimizes the waste of resources and shortens the duration of tasks. To gauge the effectiveness of the suggested framework, the article logs the experiments carried on an extensive scale which shows the superiority of the suggested framework in comparison to the traditional ones in a large-scale cloud setting
References
[1] Q. Weng, W. Xiao, Y. Yu, W. Wang, C. Wang, J. He, Y. Li, L. Zhang, W. Lin, and Y. Ding, “MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters,” in Proc. of USENIX NSDI, 2022.
[2] M. Yu, B. Ji, H. Rajan, and J. Liu, “On Scheduling Ring-All-Reduce Learning Jobs in Multi-Tenant GPU Clusters with Communication Contention,” in Proc. of ACM MobiHoc, 2022.
[3] Kathiresan, G., “Cost-Efficient and Scalable GPU Scheduling Strategies in Multi-Tenant Cloud Environments for AI Workloads,” International Journal of Computer Science and Information Technology Research, vol. 6, no. 4, pp. 1-12, 2025.
[4] E. Bampis, K. Dogeas, A. V. Kononov, G. Lucarelli, and F. Pascual, “Scheduling with Untrusted Predictions,” in Proc. of IJCAI, 2022.
[5] M. Yu, C. Wu, B. Ji, and J. Liu, “A Sum-Of-Ratios Multi-Dimensional Knapsack Decomposition for DNN Resource Scheduling,” in Proc. of IEEE INFOCOM, 2021.
[6] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-LM: Training Multi-Billion Parameter Language Models Using GPU Model Parallelism,” in Proc. of NeurIPS, 2019.
[7] S. Fan, Y. Rong, C. Meng, Z. Cao, S. Wang, Z. Zheng, C. Wu, G. Long, J. Yang, L. Xia et al., “DAPPLE: A Pipelined Data Parallel Approach for Training Large Models,” in Proc. of ACM PPoPP, 2021.
[8] D. Narayanan, A. Phanishayee, K. Shi, X. Chen, and M. Zaharia, “Memory-Efficient Pipeline-Parallel DNN Training,” in Proc. of ICML, 2021.
[9] S. K. Karmaker, M. M. Hassan, M. J. Smith, L. Xu, C. Zhai, and K. Veeramachaneni, “AutoML to Date and Beyond: Challenges and Opportunities,” ACM Computing Surveys (CSUR), vol. 54, no. 8, pp. 1–36, 2021.
[10] J. M. Tarnawski, D. Narayanan, and A. Phanishayee, “Piper: Multidimensional Planner for DNN Parallelization,” in Proc. of NeurIPS, 2021.
[11] Q. Wang, S. Shi, C. Wang, and X. Chu, “Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs,” in Proc. of IEEE INFOCOM, 2020.
[12] Y. Peng, Y. Bao, Y. Chen, C. Wu, and C. Guo, “Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters,” in Proc. of EuroSys, 2018.
[13] A. B. Faisal, N. Martin, H. M. Bashir, S. Lamelas, and F. R. Dogar, “When Will My ML Job Finish? Toward Providing Completion Time Estimates through Predictability-Centric Scheduling,” in Proc. of USENIX OSDI, 2024

