面向边缘智能感知场景的大语言模型轻量化部署技术研究
Research on Lightweight Deployment Technology of Large Language Model for Edge Intelligent Perception Scenarios
-
摘要: 随着大语言模型(Large Language Model,LLM)在自然语言处理领域取得突破性进展,将其部署于边缘智能感知场景以实现实时语义理解与智能交互已成为当前研究热点。然而,边缘计算平台固有的算力、内存及功耗约束与LLM的海量参数之间存在显著矛盾,极大地制约了LLM的规模化部署。基于此,首先,以NVIDIA Jetson Orin NX 16GB模组为硬件部署平台、ERNIE-4.5-0.3B为基准模型,建立基于Roofline的推理性能预测模型;其次,揭示小模型中词表映射层对量化误差的敏感性机理,并据此提出差异化量化策略;最后,利用离线知识蒸馏方法将大模型的逻辑推理能力迁移至小模型。实验结果表明,理论性能预测与实际评测误差小于5%,PTQ W8A16方案在1.55倍压缩比下实现30.43%的推理加速及语义精度无损,Countdown推理准确率从2.20%提升至15.40%,可为资源受限环境下的边缘语义智能感知系统构建提供有益的工程参考。Abstract: With the significant progress made by large language model(LLM)in the field of natural language processing,deploying them in edge intelligent perception scenarios to achieve real-time semantic understanding and intelligent interaction has become an important research trend.However,the inherent computational power,memory,and power consumption constraints of edge platforms,combined with the massive parameters of LLM,present a significant contradiction,severely restricting the large-scale deployment of LLM.Firstly,using the NVIDIA Jetson Orin NX 16GB module as the hardware deployment platform and ERNIE-4.5-0.3B as the benchmark model,a prediction model for inference performance based on Roofline is established.Secondly,the sensitivity mechanism of the word table mapping layer in small models to quantization errors is revealed,and a differentiated quantization strategy is proposed accordingly.Finally,the logical reasoning ability of the large model is transferred to the small model using the offline knowledge distillation method.Experimental results show that the theoretical performance prediction error is less than 5% compared to the actual evaluation,the PTQ W8A16 scheme achieves a 30.43% inference acceleration with a compression ratio of 1.55 while maintaining semantic accuracy,and the accuracy of the Countdown inference task increases from 2.20% to 15.40%,providing a beneficial engineering reference for the construction of edge semantic intelligent perception systems in resource-constrained environments.
下载: