Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WikiText
anchor
Findings
IC-1428
OPT 6.7B exhibits over 90% activation sparsity in FFN layers, reducing inference from 6.6G to 4.5G flops per token, while Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero sparsity
[eval]
IC-1429
OPT 6.7B exhibits aggregated sparsity where approximately 50% of neurons remain unused across the first 150 tokens, with a non-random reuse pattern enabling 1.27x speculative decoding speedup at gamma=16
[eval]
IC-205
Gemma 2's SAE features exhibit depth-dependent organization, with polysemantic features in early layers and persistent, matchable features in later layers
[eval]
IC-365
70B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity
[eval]
IC-367
Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behavior
[eval]
IC-425
Pretrained GPT-2 Large fails at structural in-context learning on unseen tokens in a syllogism task
[eval]
IC-916
Pythia models show scale-dependent last-layer averaging barriers: 70M exhibits a barrier of ~13 while 410M shows ~1
[eval]