Hardware–Algorithm Co-Design for Large-Scale Generative AI on Domestic AI Accelerator Architectures
Keywords:
hardware–algorithm co-design, domestic AI accelerators, generative AI, system architecture, sustainability, AI governanceAbstract
The rapid advancement of large-scale generative artificial intelligence has placed extraordinary demands on computing infrastructure, prompting a global reexamination of hardware architectures and their symbiotic relationship with algorithm design. While the dominant paradigm has long relied on a narrow set of general-purpose accelerators produced by a handful of transnational corporations, a growing body of research and industrial practice is turning toward domestically developed AI accelerators as vehicles for technological sovereignty, supply chain security, and tailored efficiency. This paper presents a comprehensive systems-level analysis of hardware–algorithm co-design strategies for training and deploying large generative models on domestic accelerator platforms. The discussion moves beyond isolated circuit-level optimizations to encompass structural trade-offs across the entire stack, including memory hierarchy, interconnect topology, numerical representation, and compiler infrastructure. Special attention is given to the unique challenges of mapping transformer-based architectures, mixture-of-experts models, and diffusion processes onto non-traditional silicon substrates that often feature distinct dataflow patterns and on-chip buffer organizations. The paper further examines the infrastructural and governance dimensions that accompany the rise of domestic accelerators, ranging from software ecosystem maturity and benchmarking standards to sustainability imperatives, algorithmic fairness, and international technology policy. By integrating architectural constraints with socio-technical considerations, the analysis reveals that co-design is not merely a technical optimization but a strategic discipline that shapes the adaptability, robustness, and democratic character of future AI systems. The conclusion articulates a forward-looking research agenda that situates domestic accelerator co-design within broader debates on innovation autonomy and responsible AI deployment.
References
1. Sze, V., Chen, Y.-H., Yang, T.-J., & Emer, J. S. (2017). Efficient processing of deep neural networks: A tutorial and survey. Proceedings of the IEEE, 105(12), 2295–2329.
2. Chen, Y.-H., Krishna, T., Emer, J. S., & Sze, V. (2017). Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks. IEEE Journal of Solid-State Circuits, 52(1), 127–138.
3. Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., ... & Yoon, D. H. (2021). Ten lessons from three generations of Tensor Processing Units. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) (pp. 1–14). IEEE.
4. Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., ... & Wu, H. (2018). Mixed precision training. In International Conference on Learning Representations (ICLR).
5. Lepikhin, D., Lee, H., Xu, Y., Chen, W., Firat, O., Huang, Y., ... & Chen, Z. (2021). GShard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations (ICLR).
6. Liao, H., Tu, J., Xia, J., & Zhou, X. (2021). DaVinci: A scalable architecture for deep neural network computing. In 2021 IEEE Hot Chips 33 Symposium (HCS) (pp. 1–24). IEEE.
7. Chen, C., Wang, C., Li, Y., Wan, Z., Geng, M., Xiao, J., ... & Peng, Y. (2026). JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators. arXiv preprint arXiv:2606.28421.
8. Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., ... & Krishnamurthy, A. (2018). TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI) (pp. 578–594).
9. Ragan-Kelley, J., Barnes, C., Adams, A., Paris, S., Durand, F., & Amarasinghe, S. (2013). Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI) (pp. 519–530).
10. Zheng, L., Jia, C., Sun, M., Wu, Z., Yu, C. H., Haj-Ali, A., ... & Gonzalez, J. E. (2020). Ansor: Generating high-performance tensor programs for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI) (pp. 863–879).
11. Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC) (pp. 1–16). IEEE.
12. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., ... & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
13. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650).
14. Shafiee, A., Nag, A., Muralimanohar, N., Balasubramonian, R., Strachan, J. P., Hu, M., ... & Srikumar, V. (2016). ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) (pp. 14–26).
15. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML) (pp. 6105–6114).
16. Khan, S. M., & Mann, A. (2020). AI chips: Technology trends and policy considerations. Center for Security and Emerging Technology, Georgetown University.
17. Meltzer, J. P. (2020). The geopolitics of artificial intelligence: Implications for international trade and security. Brookings Institution.
18. Hooker, S. (2021). Moving beyond “algorithmic bias is a data problem”. Patterns, 2(4), 100241.
19. Wan, Z., Geng, M., Chen, C., & Peng, Y. (2024). Fairness-aware quantization for neural network accelerators. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshop on Responsible AI.
20. Jordan, M. I., & Mitchell, T. M. (2015). Machine learning: Trends, perspectives, and prospects. Science, 349(6245), 255–260.
21. MLCommons. (2023). MLPerf inference benchmark results. Retrieved from https://mlcommons.org/benchmarks/inference/
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Data Intelligence and AI Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.