IC-1464Vicuna-7B and Vicuna-13B achieve only 14.2% and 18.0% average F1 on zero-shot NER, trailing ChatGPT by over 20 points

Wenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen, Hoifung Poon

SourceUniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition

Vicuna, a model fine-tuned from LLaMA using ChatGPT conversations, is evaluated zero-shot on the same UniversalNER benchmark. Vicuna-7B scores 14.2% and Vicuna-13B scores 18.0% average F1, both far below ChatGPT's 34.9%. The paper notes that LLaMA and Alpaca perform even worse, close to zero F1. In the out-of-domain evaluation (Table 3), Vicuna-7B scores 13.0% and Vicuna-13B scores 17.5% average F1.

Evidence
correlational
Key metric
Vicuna-7B: 14.2% avg F1 (zero-shot, 43 datasets); Vicuna-13B: 18.0% avg F1; out-of-domain: 13.0% (7B), 17.5% (13B)
Caveat
Vicuna was designed for general instruction following, not targeted NER; the gap with ChatGPT reflects the difference between generic distillation and targeted distillation.
Model
Vicuna
Related work
Chiang et al. 2023 (Vicuna) [context]
Related findings
IC-1463, IC-1465
Extraction
automatic-extraction