AI-Based Resource Scheduling for Cloud-Native Machine Learning Inference Platforms
Keywords:
AI-based scheduling; cloud-native platforms; machine learning inference; resource management; fairness; sustainability; robustness; governanceAbstract
Cloud-native machine learning inference platforms have become critical infrastructure for delivering low-latency, reliable, and scalable predictions across web services, mobile applications, industrial systems, and public-sector decision support. Traditional resource scheduling approaches, developed primarily for stateless microservices or batch analytics, struggle with the heterogeneous compute profiles, dynamic request patterns, model lifecycle churn, and governance requirements of modern inference workloads. This paper examines artificial intelligence based resource scheduling for cloud-native machine learning inference platforms from a system-level and socio-technical perspective. It argues that scheduling cannot be reduced to an optimization problem over central processing unit, memory, and accelerator allocation; rather, it is entangled with architectural choices, deployment topology, observability, fairness, sustainability, robustness, and policy. The analysis first situates inference platforms within cloud-native orchestration and machine learning operations, then reviews architectural constraints such as model serving graphs, autoscaling, accelerator fragmentation, and multi-tenant isolation. It subsequently evaluates AI-based scheduling methods, including reinforcement learning for placement, learned admission control, predictive autoscaling, and pipeline-aware provisioning, while emphasizing structural trade-offs between latency, cost, utilization, fairness, and resilience. The paper further considers governance and sustainability implications, particularly the risk that opaque learned schedulers may reproduce or amplify existing inequities, obscure accountability, and increase energy consumption. Through case illustrations and cross-domain comparisons, the paper identifies design principles for accountable, adaptive, and sustainable scheduling. It concludes that effective AI-based scheduling for cloud-native inference requires hybrid architectures, transparent policy constraints, robust evaluation, and continuous human oversight rather than fully autonomous optimization.
References
1. Buyya, R., Yeo, C. S., Venugopal, S., Broberg, J., & Brandic, I. (2009). Cloud computing and emerging IT platforms: Vision, hype, and reality for delivering computing as the 5th utility. Future Generation Computer Systems, 25(6), 599–616. https://doi.org/10.1016/j.future.2008.12.001
2. Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., Lee, G., Patterson, D., Rabkin, A., Stoica, I., & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58. https://doi.org/10.1145/1721654.1721672
3. Barroso, L. A., Clidaras, J., & Hölzle, U. (2013). The datacenter as a computer: An introduction to the design of warehouse-scale machines (2nd ed.). Morgan & Claypool. https://doi.org/10.2200/S00516ED2V01Y201306CAC024
4. Burns, B., Grant, B., Oppenheimer, D., Brewer, E., & Wilkes, J. (2016). Borg, Omega, and Kubernetes. Communications of the ACM, 59(5), 50–57. https://doi.org/10.1145/2890784
5. Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., & Wilkes, J. (2015). Large-scale cluster management at Google with Borg. In Proceedings of the Tenth European Conference on Computer Systems (Article 18). ACM. https://doi.org/10.1145/2741948.2741964
6. Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. https://doi.org/10.1145/2934664
7. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems 28 (pp. 2503–2511). Curran Associates.
8. Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., & Stoica, I. (2017). Clipper: A low-latency online prediction serving system. In 14th USENIX Symposium on Networked Systems Design and Implementation (pp. 613–627). USENIX Association.
9. Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., & Zhang, Z. (2015). MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv. https://arxiv.org/abs/1512.01274
10. Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., ... Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (pp. 265–283). USENIX Association.
11. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., ... Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32 (pp. 8024–8035). Curran Associates.
12. Reddi, V. J., Cheng, C., Kanter, D., Mattson, P., Schmuelling, G., Wu, C.-J., Anderson, B., Breughe, M., Charlebois, M., Chou, W., Chukka, R., Coleman, C., Davis, S., Deng, C., Diamos, G., Duke, J., Fink, D., Fung, P., Gopalakrishnan, K., ... Zhou, Y. (2020). MLPerf inference benchmark. In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (pp. 446–459). IEEE. https://doi.org/10.1109/ISCA45697.2020.00045
13. Mattson, P., Reddi, V. J., Cheng, C., Coleman, C., Diamos, G., Kanter, D., Micikevicius, P., Patterson, D., Schmuelling, G., Tang, H., Wei, G.-Y., & Wu, C.-J. (2020). MLPerf: An industry standard benchmark suite for machine learning performance. IEEE Micro, 40(2), 8–16. https://doi.org/10.1109/MM.2020.2974843
14. Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., & Su, B.-Y. (2014). Scaling distributed machine learning with the parameter server. In 11th USENIX Symposium on Operating Systems Design and Implementation (pp. 583–598). USENIX Association.
15. Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., & Ng, A. Y. (2012). Large scale distributed deep networks. In Advances in Neural Information Processing Systems 25 (pp. 1223–1231). Curran Associates.
16. Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., & Zaharia, M. (2019). PipeDream: Generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (pp. 1–15). ACM. https://doi.org/10.1145/3341301.3359646
17. Crankshaw, D., Sela, O., Mo, X., Zumar, C., Stoica, I., Gonzalez, J. E., & Tumanov, A. (2020). InferLine: Latency-aware provisioning and scheduling for machine learning prediction pipelines. In Proceedings of the 11th ACM Symposium on Cloud Computing (pp. 105–119). ACM. https://doi.org/10.1145/3419111.3421294
18. Mao, H., Schwarzkopf, M., Venkatakrishnan, S. B., Meng, Z., & Alizadeh, M. (2019). Learning scheduling algorithms for data processing clusters. In Proceedings of the ACM Special Interest Group on Data Communication (pp. 270–288). ACM. https://doi.org/10.1145/3341302.3342080
19. Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., Norouzi, M., Bengio, S., & Dean, J. (2017). Device placement optimization with reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning (pp. 2430–2439). PMLR.
20. Chen, L., Lingys, J., Chen, K., & Liu, F. (2018). Auto: Scaling deep reinforcement learning for datacenter-scale automatic traffic optimization. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication (pp. 191–205). ACM. https://doi.org/10.1145/3230543.3230551
21. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63. https://doi.org/10.1145/3381831
22. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1355
23. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). ACM. https://doi.org/10.1145/3442188.3445922
24. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D'Oliveira, R. G. L., Eichner, H., El Rouayheb, S., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., ... Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210. https://doi.org/10.1561/2200000083
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Computational Intelligence Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.