IC-709Swin Transformer parameter redundancy is depth-dependent, with higher layers retaining fewer parameters than lower layers under module-aware pruning

Yang He, Joey Tianyi Zhou

SourceData-independent Module-aware Pruning for Hierarchical Vision Transformers

The authors analyse which weights in Swin-S are most important using their module-level distortion metric and observe a clear depth-dependent pattern in the ratio of kept parameters. Lower layers (early in the network) retain a larger fraction of their parameters, while higher layers (deeper in the network) are pruned more aggressively, indicating greater redundancy in deeper layers. Additionally, within the first and second Swin blocks, attention-related layers (att-qkv, att-prj) retain a higher ratio of parameters than MLP-related layers (mlp-fc1, mlp-fc2), suggesting attention is more critical for low-level feature extraction. This depth-dependent redundancy structure is what makes uniform magnitude pruning suboptimal, as it does not account for the varying importance across depths.

Evidence
observational
Key metric
Swin-S pruned 45% of parameters: lower layers keep a larger ratio of parameters than higher layers (Figure 5); in first and second blocks, attention layers (1,2,5,6) have larger kept ratios than MLP layers (3,4,7,8); Swin-B top-5 drop 0.07% at 52.5% params removed; Swin-S top-5 drop -0.03% at 33.2% params removed
Caveat
The depth-dependent pattern is observed using the authors' own module-level importance metric; a different importance criterion might reveal a different depth profile. The analysis is limited to Swin Transformers and may not generalise to other hierarchical ViT architectures.
Model
Swin Transformer Swin-S, Swin-B, Swin-T
Concepts
Depth-dependent structure
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Methods
Uniform Magnitude Pruning [compared-to]
Related work
Swin Transformer [builds-on]
Extraction
automatic-extraction