The paper examines the US fairness metric across multiple sizes within six model families (Pythia 70M-12B, OPT 125M-6.7B, LLaMA-2 7B-70B, MPT 7B-30B, OLMo 1B-7B, Mistral/Mixtral 7B-8x7B). No consistent monotonic trend is found: within each family, fairness scores do not reliably increase or decrease with model size. For example, in Pythia on USE-5, the 410M model (28.67%) outperforms the 1.4B (18.37%) and 2.8B (18.23%), while the 12B (31.33%) is highest. The authors conclude that model size alone does not determine gender fairness in non-stereotypical settings.