SelectInfer introduces a neuron-level optimization framework for running LLMs on resource-constrained edge devices. An offline profiler identifies task-specific and general-purpose neurons. During inference, selective loading keeps only important neurons in memory, while selective computation dynamically evaluates the most relevant neurons. The authors report evaluations across multiple datasets showing reduced memory footprint and computation with preserved task performance. However, the supplied abstract does not provide model names, dataset names, quantitative speed or memory reductions, accuracy deltas, hardware details, or comparisons against pruning and quantization baselines.
No heat snapshots are available in the last 24 hours.