IC-1200GPT-2-XL's untruncated next-token log-probability matrix has rank saturating at its hidden dimensionality of 1600, while truncation sampling produces post-truncation distributions whose estimated rank grows far beyond 1600

Matthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta, Ashish Sabharwal

SourceClosing the Curious Case of Neural Text Degeneration

The authors run GPT-2-XL on OpenWebText prefixes, concatenate the conditional log-distributions, and compute the rank of the resulting matrix. Without truncation (ancestral sampling), the rank saturates at 1600, matching the model's hidden dimensionality as predicted by the softmax bottleneck theory. With nucleus, η, or ε truncation, the estimated rank continues to grow with the number of prefixes, reaching well past 1600 (up to ~5000 in the figure). This confirms that truncation sampling effectively breaks the low-rank constraint on the model's output space.

Evidence
correlational
Key metric
GPT-2-XL hidden dimensionality 1600; untruncated rank saturates at 1600; truncated (nucleus, η, ε) rank estimate grows to ~5000 at 10000 prefixes
Caveat
Limited by the number of prefixes that can be processed (matrix is 50257-dimensional per prefix); rank is estimated, not exact; only GPT-2-XL tested
Model
GPT-2
Datasets
OpenWebText / OpenWebText-10k [eval]
Methods
Nucleus Sampling [compared-to]
Related work
Breaking the softmax bottleneck (Yang et al. 2018) [builds-on]
Related findings
IC-1199
Extraction
automatic-extraction