IC-046Tulu-2-13B exhibits gender bias in both its internal binding representation and its outputs, with the output-level bias being stronger than the representation-level bias
In a gender-bias evaluation with 400 templated contexts specifying genders and stereotypically male or female occupations, both prompting and probing (binding similarity) show higher accuracy on pro-stereotypical than anti-stereotypical contexts. However, the probing method is significantly less biased than prompting. The authors conclude that gender bias influences the model in at least two ways: it affects how binding is done (internal world state construction) and how the model responds to queries about that state. Probing mitigates the latter but not the former.
Evidence
correlational
Caveat
The evaluation uses a small closed world of 14 occupations from Winobias and 2 genders; the calibrated accuracy (controlling for label bias) does not significantly change the results.