Engineering AI Factories: Scalable HPC Design Patterns for Next-Generation Foundation Model Infrastructure

Authors

  • Rakesh Challa Principal Engineer, Dell Technologies, USA Author

DOI:

https://doi.org/10.15680/IJCTECE.2023.0604013

Keywords:

AI Factories, GPU Clusters, High Performance Computing, RDMA, Network Optimization, InfiniBand, NCCL, Distributed Training

Abstract

The paper examines the design of AI Factories, which are massive HPC systems constructed to provide training and serve the current AI models. The analysis is dedicated to the scaling of systems of thousands of GPUs, enhancing network performance, automatic deployment, and evaluation of the results with the help of popular benchmarks. The findings indicate that there can be a high level of scalability of the system and that the efficiency of the system declines from 92% to 82.5% when using 512 and 8192 GPUs respectively. Tests of networks indicate that InfiniBand has decreased the latency delay from 12 µs to 6 µs and work turnaround period had been decreased from 145 minutes to 120 minutes in comparison to Ethernet. Automation has raised the success rate from 91% to 98% through deployment automation lowering failures on the systems. The benchmark performance is that HPL efficiency remains in the range of 78%-90%, and NCCL bandwidth increases to 210 GB/s. These findings affirm that large scale AI workloads can be efficiently enabled by patterns of designs optimized. The paper shows the feasible ways to minimize failures in deployments and enhance training performance. The paper gives a complete roadmap on how scalable AI infrastructure can be established to provide support to millions of users and the ongoing training of the AI models.

 

References

[1] W. Jiang, F. Ren, and J. Wang, “Survey on link layer congestion management of lossless switching fabric,” Computer Standards & Interfaces, vol. 57, pp. 31–35, Nov. 2017, doi: 10.1016/j.csi.2017.11.002.

[2] Y. Zhu et al., “Congestion control for Large-Scale RDMA deployments,” 2015. [Online]. Available: https://conferences.sigcomm.org/sigcomm/2015/pdf/papers/p523.pdf

[3] J. Xue, M. U. Chaudhry, B. Vamanan, T. N. Vijaykumar, and M. Thottethodi, “DART: divide and specialize for fast response to congestion in RDMA-based datacenter networks,” arXiv.org, May 28, 2018. https://arxiv.org/abs/1805.11158

[4] M. Hemmatpour, B. Montrucchio, and M. Rebaudengo, “Communicating Efficiently on Cluster-Based Remote Direct Memory Access (RDMA) over InfiniBand Protocol,” Applied Sciences, vol. 8, no. 11, p. 2034, Oct. 2018, doi: 10.3390/app8112034.

[5] S. Novakovic et al., “Storm: a fast transactional dataplane for remote data structures,” arXiv (Cornell University), Feb. 2019, doi: 10.48550/arxiv.1902.02411.

[6] X. Liu, Y. Hua, X. Li, and Q. Liu, “Write-Optimized and consistent RDMA-based NVM systems,” arXiv (Cornell University), Jun. 2019, doi: 10.48550/arxiv.1906.08173.

[7] R. Mittal et al., “TIMELY: RTT-based congestion control for the datacenter,” 2015. [Online]. Available: https://web.stanford.edu/class/cs244/papers/timely-sigcomm2015.pdf

[8] R. Mittal et al., “TIMELY,” TIMELY, pp. 537–550, Aug. 2015, doi: 10.1145/2785956.2787510.

[9] J. Xue, M. U. Chaudhry, B. Vamanan, T. N. Vijaykumar, and M. Thottethodi, “DART: divide and specialize for fast response to congestion in RDMA-based datacenter networks,” arXiv.org, May 28, 2018. https://arxiv.org/abs/1805.11158

[10] V. Addanki, O. Michel, and S. Schmid, “PowerTCP: Pushing the performance limits of datacenter networks,” arXiv (Cornell University), Dec. 2021, doi: 10.48550/arxiv.2112.14309.

[11] Y. Zhu et al., “Congestion control for Large-Scale RDMA deployments,” ACM SIGCOMM Computer Communication Review, vol. 45, no. 4, pp. 523–536, Aug. 2015, doi: 10.1145/2829988.2787484.

[12] R. Mittal et al., “Revisiting network support for RDMA,” Revisiting Network Support for RDMA, pp. 313–326, Aug. 2018, doi: 10.1145/3230543.3230557.

[13] C. Tessler et al., “Reinforcement learning for datacenter congestion control,” arXiv (Cornell University), Feb. 2021, doi: 10.48550/arxiv.2102.09337.

[14] S. Lee, Y. Kim, H. Woo, and I. Yeom, “Efficient User-Level Multi-Path utilization in RDMA networks,” IEEE Access, vol. 9, pp. 127619–127629, Jan. 2021, doi: 10.1109/access.2021.3110840.

[15] F. Tian, W. Feng, Y. Zhang, and Z.-L. Zhang, “A novel software-based multi-path RDMA solutionfor data center networks,” arXiv (Cornell University), Sep. 2020, doi: 10.48550/arxiv.2009.00243.

[16] S. Liu et al., “NetReduce: RDMA-Compatible In-Network reduction for distributed DNN training acceleration,” arXiv (Cornell University), Sep. 2020, doi: 10.48550/arxiv.2009.09736.

[17] S. Jha et al., “A study of network congestion in two supercomputing High-Speed interconnects,” arXiv.org, Jul. 11, 2019. https://arxiv.org/abs/1907.05312

[18] Q. Li et al., “From RDMA to RDCA: toward High-Speed last mile of data center networks using remote direct cache access,” arXiv (Cornell University), Nov. 2022, doi: 10.48550/arxiv.2211.05975.

[19] T. Ziegler, V. Leis, and C. Binnig, “RDMA communciation Patterns,” Datenbank-Spektrum, vol. 20, no. 3, pp. 199–210, Sep. 2020, doi: 10.1007/s13222-020-00355-7.

[20] F. Mizero, M. Veeraraghavan, Q. Liu, R. D. Russell, and J. M. Dennis, “A dynamic congestion management system for InfiniBand networks,” Supercomputing Frontiers and Innovations, vol. 3, no. 2, Sep. 2016, doi: 10.14529/jsfi160201.

[21] S. Zhou et al., “SR-DCQCN: Combining SACK and ECN for RDMA Congestion Control,” SR-DCQCN: Combining SACK and ECN for RDMA Congestion Control, pp. 788–794, Dec. 2022, doi: 10.1109/iccc56324.2022.10065950.

[22] S. Ma, J. Jiang, W. Wang, and B. Li, “Fairness of Congestion-Based Congestion Control: Experimental Evaluation and Analysis,” arXiv (Cornell University), Jun. 2017, doi: 10.48550/arxiv.1706.09115.

[23] M. K. Aguilera, N. Ben-David, R. Guerraoui, V. Marathe, and I. Zablotchi, “The impact of RDMA on agreement,” arXiv.org, May 29, 2019. https://arxiv.org/abs/1905.12143

[24] W. Roediger, T. Muehlbauer, A. Kemper, and T. Neumann, “High-Speed Query Processing over High-Speed Networks,” arXiv (Cornell University), Feb. 2015, doi: 10.48550/arxiv.1502.07169.

[25] T. Khan, S. Rashidi, S. Sridharan, P. Shurpali, A. Akella, and T. Krishna, “Impact of ROCE congestion Control Policies on distributed training of DNNs,” arXiv (Cornell University), Jul. 2022, doi: 10.48550/arxiv.2207.10898.

Downloads

Published

2023-07-18

How to Cite

Engineering AI Factories: Scalable HPC Design Patterns for Next-Generation Foundation Model Infrastructure. (2023). International Journal of Computer Technology and Electronics Communication, 6(4), 7342-7351. https://doi.org/10.15680/IJCTECE.2023.0604013