Scholay

学术搜索 · AI 审稿 · LaTeX 协作

An Adaptive and Scalable Framework for Resource-Efficient Deployment of Mixture of Experts in LLM-Based Intelligent IoT Networks

作者:C.M. Liu, Yangfan Li, Cen Chen, Hailan Kuang, Xiaolin Ma, Xiaofeng Zou, Jing Liu, Zeyu Lu, Zhaoyuan Zhang, Xinhua Liu · 发表于:IEEE Internet of Things Journal · 年份:2025 · DOI:10.1109/jiot.2025.3550841 · 被引用次数:5 · 研究领域:IoT and Edge/Fog Computing、Distributed and Parallel Computing Systems

The exponential growth of the Internet of Things (IoT) necessitates the deployment of large-scale models capable of processing the complex and diverse data generated by IoT devices. However, the substantial memory requirements of these models pose significant challenges, especially in scenarios where rapid decision-making and low-latency responses are critical. To address these challenges, we propose three innovative strategies for optimizing large model usage in IoT environments. The first strategy is an adaptive loading scheme, which enables dynamic loading of individual model experts. The second strategy involves an expert-by-expert loading approach, further enhancing the ability to load experts as needed, which optimizes memory usage and accelerates computations. The third strategy employs an interlayer expert reuse mechanism, facilitating the efficient reuse of experts across different layers, thus enhancing response rates without compromising model accuracy. Importantly, these strategies can be directly applied to Mixture of Experts (MoE) large language models without requiring additional training, thereby providing a seamless and efficient solution for leveraging these models in memory-constrained, high-performance IoT environments.