Adaptive Model Compression and Dynamic Inference for Large-Scale Generative AI in Industrial IoT
Keywords:
adaptive model compression, dynamic inference, generative AI, Industrial IoT, edge computing, system governance, sustainability, model deploymentAbstract
The integration of large-scale generative artificial intelligence into industrial Internet of Things environments introduces a fundamental tension between model expressiveness and the operational constraints of distributed edge infrastructure. Industrial systems increasingly demand generative capabilities for anomaly synthesis, predictive maintenance, visual inspection, and human-machine interaction, yet the computational, memory, energy, and communication requirements of foundation models often exceed the capacity of field-deployed devices. This paper examines adaptive model compression and dynamic inference as systemic strategies for reconciling these competing pressures. Rather than treating compression as a one-time model reduction task, the paper argues that industrial deployment requires continuous, context-aware adaptation across heterogeneous edge-cloud hierarchies. It develops a system-level perspective that considers architectural trade-offs, runtime reconfiguration, data governance, fairness, robustness, sustainability, and regulatory compliance. The discussion draws on cross-domain cases from manufacturing, energy systems, logistics, and smart infrastructure to illustrate how adaptive compression interacts with organizational and technical constraints. The paper further addresses policy implications for industrial AI governance, including the need for auditable adaptation mechanisms, carbon-aware inference scheduling, and accountable resource allocation. It concludes that adaptive model compression should be understood not merely as an efficiency technique but as a governance infrastructure that shapes the reliability, equity, and sustainability of industrial generative AI systems.
References
1. Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. International Conference on Learning Representations.
2. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
3. Teerapittayanon, S., McDanel, B., & Kung, H. T. (2017). BranchyNet: Fast inference via early exiting from deep neural networks. 2017 26th International Conference on Computer Communications and Networks (ICCCN).
4. Laskaridis, S., Venieris, S. I., Almeida, M., Leontiadis, I., & Lane, N. D. (2020). SPINN: Synergistic progressive inference of neural networks over device and cloud. Proceedings of the 26th Annual International Conference on Mobile Computing and Networking.
5. Cheng, Y., Wang, D., Zhou, P., & Zhang, T. (2018). A survey of model compression and acceleration for deep neural networks. IEEE Signal Processing Magazine, 35(1), 126-136.
6. Konecny, J., McMahan, H. B., Yu, F. X., Richtarik, P., Suresh, A. T., & Bacon, D. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.
7. Chen, Ce, et al. "JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators." arXiv preprint arXiv:2606.28421 (2026).
8. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
9. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10684-10695.
10. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.
11. Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2019). Edge intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE, 107(8), 1738-1762.
12. Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637-646.
13. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
14. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645-3650.
15. Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., ... Hazelwood, K. (2022). Sustainable AI: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4, 795-813.
16. Barocas, S., Hardt, M., & Narayanan, A. (2019). Fairness and machine learning. fairmlbook.org.
17. Vepakomma, P., Swedish, T., Raskar, R., Gupta, O., & Dubey, A. (2018). No peek: A survey of private distributed deep learning. arXiv preprint arXiv:1812.03288.
18. Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. International Conference on Learning Representations.
19. Isard, M., Budiu, M., Yu, Y., Birrell, A., & Fetterly, D. (2007). Dryad: Distributed data-parallel programs from sequential building blocks. Proceedings of the 2nd ACM SIGOPS/EuroSys European Conference on Computer Systems, 59-72.
20. Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 248-255.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Data Intelligence and AI Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.