This paper adapts Adaptive Feature Retention (AFR), originally developed for unstructured pruning, to structured pruning of large language models. It identifies three issues: heterogeneous pruning-score distributions, loss of sign information that reflects optimization-direction consistency, and sensitivity to outliers. The proposed method combines nonlinear power transformation, sign-preserving score aggregation, and percentile-based outlier removal. Experiments on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B reportedly show accuracy comparable to unstructured pruning while delivering practical inference speedups from structured sparsity. The abstract does not provide detailed speed, memory, dataset, or ablation results.
No heat snapshots are available in the last 24 hours.