Domain Large Language Models, Model Compression, Efficient Fine-Tuning
Task Literature Survey
Work Surveys joint compression and fine-tuning methods for LLMs, including pruning, quantization, knowledge distillation, low-rank approximation, PEFT, and alignment tuning, while organizing representative toolchains and open challenges for efficient deployment under memory, latency, and hardware constraints
Reference Survey on Joint Compression and Fine-Tuning of Large Language Models: Methods, Toolchains, and Open Challenges under Resource Constraints, Neurocomputing, 2026