TY - GEN
T1 - Task-Aware Parameter-Efficient Fine-Tuning of Large Pre-Trained Models at the Edge
AU - Hu, Senkang
AU - Ma, Yanan
AU - Tao, Yihang
AU - Fang, Zhengru
AU - Fang, Zihan
AU - Deng, Yiqin
AU - Wu Kwong, Sam Tak
AU - Fang, Yuguang
PY - 2025/12
Y1 - 2025/12
N2 - Large language models (LLMs) have achieved remarkable success in various tasks, such as decision-making, reasoning, and question answering. They have been widely used in edge devices. However, fine-tuning LLMs to specific tasks at the edge is challenging due to the high computational cost and the limited storage and energy resources at the edge. To address this issue, we propose TaskEdge, a task-aware parameter-efficient fine-tuning framework at the edge, which allocates the most effective parameters to the target task and only updates the task-specific parameters. Specifically, we first design a parameter importance calculation criterion that incorporates both weights and input activations into the computation of weight importance. Then, we propose a model-agnostic task-specific parameter allocation algorithm to ensure that task-specific parameters are distributed evenly across the model, rather than being concentrated in specific regions. In doing so, TaskEdge can significantly reduce the computational cost and memory usage while maintaining performance on the target downstream tasks by updating less than 0.1% of the parameters. In addition, TaskEdge can be easily integrated with structured sparsity to enable acceleration by NVIDIA's specialized sparse tensor cores, and it can be seamlessly integrated with LoRA to enable efficient sparse low-rank adaptation. Extensive experiments on various tasks demonstrate the effectiveness of TaskEdge. © 2025 IEEE.
AB - Large language models (LLMs) have achieved remarkable success in various tasks, such as decision-making, reasoning, and question answering. They have been widely used in edge devices. However, fine-tuning LLMs to specific tasks at the edge is challenging due to the high computational cost and the limited storage and energy resources at the edge. To address this issue, we propose TaskEdge, a task-aware parameter-efficient fine-tuning framework at the edge, which allocates the most effective parameters to the target task and only updates the task-specific parameters. Specifically, we first design a parameter importance calculation criterion that incorporates both weights and input activations into the computation of weight importance. Then, we propose a model-agnostic task-specific parameter allocation algorithm to ensure that task-specific parameters are distributed evenly across the model, rather than being concentrated in specific regions. In doing so, TaskEdge can significantly reduce the computational cost and memory usage while maintaining performance on the target downstream tasks by updating less than 0.1% of the parameters. In addition, TaskEdge can be easily integrated with structured sparsity to enable acceleration by NVIDIA's specialized sparse tensor cores, and it can be seamlessly integrated with LoRA to enable efficient sparse low-rank adaptation. Extensive experiments on various tasks demonstrate the effectiveness of TaskEdge. © 2025 IEEE.
KW - Edge Computing
KW - Large Language Models
KW - Large Pre-Trained Models
KW - Parameter-Efficient Fine-Tuning
UR - http://www.scopus.com/inward/record.url?scp=105036314356&partnerID=8YFLogxK
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-105036314356&origin=recordpage
U2 - 10.1109/GLOBECOM59602.2025.11432128
DO - 10.1109/GLOBECOM59602.2025.11432128
M3 - RGC 32 - Refereed conference paper (with host publication)
T3 - Proceedings - IEEE Global Communications Conference, GLOBECOM
SP - 1035
EP - 1040
BT - GLOBECOM 2025 - 2025 IEEE Global Communications Conference
PB - IEEE
T2 - 2025 IEEE Global Communications Conference (GLOBECOM 2025)
Y2 - 8 December 2025 through 12 December 2025
ER -