
Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT
Researchers propose a cascaded pruning framework that systematically compresses large language models for edge deployment in industrial IoT environments by removing layers, attention heads, and feed-forward channels in staged phases with intermediate low-rank recovery. The work addresses a critical bottleneck in on-device inference: existing one-shot pruning methods fail catastrophically at extreme compression ratios needed for resource-constrained hardware. By formalizing the Structural Independence Assumption as a predictability condition, the authors provide a principled method to determine when per-component pruning criteria remain reliable across different architectures, potentially unlocking practical LLM deployment in manufacturing, logistics, and other industrial settings where cloud connectivity is unavailable or latency-prohibitive.62




























