Light Dark text · generative · anchor
Open-weight language model family cited in the corpus at 7B and 72B, including the instruction-tuned 7B checkpoint.
Note anchor found by search and checked against this entry's own description before it was recorded: "Qwen2 Technical Report", covering 0.5B to 72B, which spans the 7B and 72B carried here Variants Qwen 2 7B , Qwen 2 7B Instruct , Qwen 2 72B , Qwen2-72B-Instruct , BTRM_Qwen2_7B , Qwen2-0.5B-Instruct , Qwen2-1.5B , Qwen2-32B-Instruct Findings IC-007 Most LLMs do not align closely with human moral preferences on multilingual trolley problems IC-009 LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages IC-011 Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences IC-029 Large language models show conformity to group answers in multi-agent interactions IC-030 Larger language models exhibit higher independence rates and lower conformity under some protocols IC-031 Empowered persona prompts and reflection mechanisms reduce conformity in large language models IC-073 Released LLMs (GPT-4o, Llama-3.1-70B, Qwen2-7B, etc.) show limited workflow orchestration capability that degrades as workflow complexity increases IC-112 Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increases IC-113 Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputs IC-114 Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl range IC-115 NUPA performance is largely independent of model size within a family: GPT-4o and GPT-4o-mini show nearly identical performance, as do Qwen2-72B and Qwen2-7B IC-169 BT-based, DPO-based reward models, and GPT-4 as judge all exhibit significant length bias, with their scores correlating with output length rather than quality IC-188 LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categories IC-189 Most LLMs show degraded instruction satisfaction when user instructions conflict with system messages, indicating difficulty in prioritizing system message constraints IC-190 LLMs show progressive degradation in system message constraint following across multi-turn conversations, with dependent conversations degrading faster than parallel ones IC-191 Attention allocated to system messages correlates with following ability, and models do not strictly distinguish system from user messages based on marker tokens IC-213 Qwen2-0.5B-Instruct uses English-specific past tense heads and late FFN layers for morphological marking that is absent in Chinese IC-221 GPT-4o ReAct success rate drops from 47% on synchronous to 11% on asynchronous planning tasks, and all other tested LLMs show equal or worse performance IC-351 GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarks IC-352 RAG with sufficient retrieved tokens outperforms direct long-context for Qwen2-72b-instruct on >100k tasks, while at 32k the default RAG setting underperforms direct long-context for GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct IC-427 Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant data IC-436 API selection accuracy of 10 LLM-based agents degrades sharply as task complexity increases, with open-source models ≥70B matching closed-source on simpler tasks but lagging on the most complex IC-437 Extracting parameters from user queries is harder for LLM-based agents than using outputs from previous actions, and less intelligent LLMs show steeper parameter-filling degradation with task difficulty IC-438 All 10 LLM-based agents perform poorly at recognizing when they need to request input from the system or user, with overall accuracy between 30.55% and 55.18% IC-485 LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from context IC-510 Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parameters IC-549 All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench IC-550 Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B models IC-552 GPT-4, Llama-3.1-8B, and Qwen-2-72B all improve on ALFWorld and WebShop when given a generated workflow as structured prior knowledge IC-575 Four released LLMs (LLaMA-3.1-8B, Mistral-7B, Qwen2-7B, Yi-1.5-9B) can perform in-context learning on continuous vector representations projected into their embedding space, matching or outperforming few-shot ICL across text, time-series, graph, and fMRI tasks IC-576 For 10-digit numerical function regression, vector-ICL consistently outperforms few-shot ICL with raw number inputs across all four LLMs because continuous representations avoid multi-token splitting IC-601 Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases IC-602 Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows Shared mechanisms Depth-dependent structure also in Baichuan 2 , BERT , BLIP-2 , BLOOM , Chameleon , CLIP / CLIP-ViT (LC) , DeepFloyd IF , DeiT-III , DINO , DINOv2 , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-J , GPT-NeoX-20B , Griffin , I3D , Idefics , InstructBLIP , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Llama 3.2 , Llama-3.2-3B , LLaVA , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Phi , MAE , MAE-B/16 , Mamba , MiniGPT-4 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , mPLUG-Owl , MPT , MultiBERTs , MViT V2 , OLMo / OLMo base , OpenCLIP , OPT , Phi-2 , Pythia , Qwen2-VL , Qwen2.5 , RoBERTa / RoBERTa-L , RWKV , SALMONN , SAM , SlowFast , Stable Diffusion , Swin Transformer , TerraMind , TimesFormer , TSM , Uniformer , Vicuna , VideoMAE , ViT , X3D , Yi Failure mode also in AASIST , ADM , Aegis-Guard-Defensive , Alpaca , AnyLoc , AutoTikZ / DataTikZ , Baichuan , Baichuan 2 , Baichuan2-13B , BakLLaVA , Bard , BEiT , BERT , BingChat , BLIP , BLIP-2 , BLOOM , BSRNet , CF2 , Chat-UniVi-7B , ChatGLM-6B / ChatGLM-6b-2 , ChatGLM2 , ChatGPT , CLAP , Claude 1.3 , Claude 2.0 , Claude 2.1 , Claude 3 , Claude 3.5 , CLEAR , CLIP / CLIP-ViT (LC) , CLIP4Clip , CLIPBERT , CLIPCap , CLMBR-T-BASE , CloFNet , Code Llama , CodeGeex2 , CodeGen , CodeLlama-13B , CodeLlama-34B , CogVLM2 , Cohere Command R , CoMEt , Command R+ , CONCH , CycleGAN , DALL-E , DALL·E 2 , DALL·E 3 , DASR , DECAF , DeepSeek-2-Chat , DeepSeek-2-Coder , DeepSeek-V2-0628 , DeepSeek-VL , DeepSeek-VL2 , DeiT , DeiT-III , Depth Anything , DETR-R50 , DimeNet++ , DINO , DINO-FocalNet-Large , DINOv2 , EGNN , Emu2 , EquiformerV2 , ESCN , ESM-2 , ESM3 , ESRGAN , EVA-CLIP , EVE , Falcon , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 , FLAN-T5 , Florence-2 , FLUX / FLUX1 , FoveaBox , Fuyu , Galactica-6.7B , GAT , 3D Gaussian Splatting , GCN , Gemini , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GIN , GLIDE , GLM-4 , GLM-4V , GloVe , GP-UNIT , GPT-2 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4.1 , GPT-4o , GPT-J , GPT-NeoX-20B , GraphSAGE , Grounding DINO , Guanaco , GVP , Hawkeye , HiFaceGAN , HPS , HuggingGPT , IDDPM , Idefics , Idefics2 , ImageBind , ImageBind-LLM-7B , Imagen Video , ImageReward , 12-in-1 , InstructBLIP , InstructGPT , InternLM-2.5-7B , InternLM-XComposer2-VL , InternVideo , InternVL-1.5 , InternVL2 , Koala , LegalBERT , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Llama 3.2 , Llama-3.2-3B , Llama-3-2-Vision , LLaMA-Adapter v2 , Llama Guard , Llama Guard 2 , Llama-Guard 3 , Llama-VID , LLaVA , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Med , LLaVA-NeXT / LLaVA 1.6 , LLaVA-OneVision , LongVA-7B , LOVT , LWM-1M-JAX , MACE , MAE , Med-Flamingo , Merlot Reserve , MGCA , MiDaS , MiniCPM-V , MiniGPT-4 , Mip-Splatting , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mistral Large 2 , Mistral Large V2 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , MobileNetV2 , Molmo , MolmoE-7B , Momentor , Moondream2 , Moonshot-v1-8k , mPLUG-2 , mPLUG-Owl , mPLUG-Owl3 , mPLUG-Owl2 , MPT , MSA Transformer , MultiBERTs , Nova Canvas , Nova Lite , Nova Pro , O1 / OpenAI-o1-preview , O3 , O4-mini , OLMo / OLMo base , OneLLM , OpenAI Moderation , OpenChat-3.5-0106 , OpenCLIP , OpenFlamingo , OPT , Otter , Otter-7B , PaLM 2 , PaLM 62B , PandaGPT-7B , PerSAM , Phi-3 , Phi-3.5 Mini Instruct , PickScore , PLIP , Prismatic , ProGen-2 , Pythia , Qwen1.5 , Qwen 2.5 72B Instruct , Qwen2-VL , Qwen-Audio , Qwen-VL , Qwen2.5 , Qwen2-Audio , R2D2 , RadFM , RCExplainer , RD-SR , RealESRNet , Reprover , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , RetinaNet-R50 , RivaGAN , RS-LDS , RT-1 , RT-2 , SALMONN , SAM , SAM 2 , SAULLM 54B , Scaffold-GS , SchNet , Seed-LLaMA-8B , SGC , SIREN , Sketch Transformer , SLD-max , SLD-medium , SLD-strong , SLDS , SLIMG , SpeechGPT , SphereNet , SRResNet , Stable Diffusion , StackLLaMA , Starcoder , StegaStamp , StyleGAN2-ADA , Swin Transformer , T5 , TD-MPC , TerraMind , TimeChat , TranceptionEVE , TreeRing , Tulu 2 , UnifiedQA , UniPerceiver , UNITER , UniVL , Van Gogh Diffusion , VERA , VGG / VGG13 , Vicuna , Video-Chat-7B , Video-ChatGPT , Video-LLaMA , Video-LLaMA-2-13B , Video-LLaVA , VideoCLIP , ViLA-8B , ViLBERT , VindLU , VioLET , ViRTex , ViT , ViV1T , VTG-LLM , WildGuard , Wizardlm , X-CLIP , X-InstructBLIP-7B , XGen-MM , Xlm-R , Zephyr-7B-beta Method artefact also in Baichuan , CLIP / CLIP-ViT (LC) , ConvNeXt , EfficientNet , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-3.5 / ChatGPT-3.5 , GPT-4o , GPT-J , InternLM , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , MAP-NEO , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , OLMo / OLMo base , OpenCLIP , OpenLLaMA , OPT , Pythia , Qwen1.5 , Qwen2.5 , RedPajama , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , Skywork , Stable Diffusion , StableLM , TerraMind , ViT , Yi , ZiYA2 Positional bias also in BERT , ChatGPT , Claude 3 , Claude 3.5 , Falcon , Fuyu , Gemini , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GLIDE , GPT-2 , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4.1 , GPT-4o , GPT-J , InstructGPT , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , MPT , O1 / OpenAI-o1-preview , O3 , O4-mini , PaLM 2 , Phi-3 , Pythia , Qwen1.5 , Qwen 2.5 72B Instruct , Stable Diffusion , Sybil , Vicuna Scale-dependent behaviour also in Aquila-7B , BEiT , BERT , BLOOM , Claude 2.1 , Claude 3 , Claude 3.5 , CLIP / CLIP-ViT (LC) , Code Llama , CodeGen , Cohere Command R , DeepSeek LLM , DeepSeekMoE , DeiT-III , DINO , DINOv2 , EquiformerV2 , ESCN , Falcon , FLAN-T5 , Gemini 1.0 Pro , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GPT-2 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-J , GPT-Neo , I3D , Idefics , InternLM-2.5-7B , InternLM-XComposer2-VL , InternLM2 , InternVL-1.5 , InternVL2 , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Llama 3.2 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 , LongVA-7B , MAE , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Moirai , MPT , MViT V2 , O1 / OpenAI-o1-preview , OLMo / OLMo base , OpenCLIP , OpenFlamingo , OpenLLaMA , OPT , PaLM 2 , Phi-3 , Platypus2-Instruct-70B , Pythia , Qwen , Qwen1.5 , Qwen 2.5 72B Instruct , Qwen2-VL , Qwen-Audio , Qwen2.5 , RedPajama-INCITE , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SlowFast , Solar 10.7B , Stable Diffusion , StableLM , Swin Transformer , T5 , TerraMind , text-ada-001 , TigerBot-70B-Chat , TimesFormer , TSM , Tulu 2 , Uniformer , Vicuna , VideoMAE , ViLA-8B , Wizardlm , X3D , XGLM , Xwin-LM-70B-v0.1 , Yi Shortcut also in BakLLaVA , BLIP-2 , Claude 3 , Claude 3.5 , CLIP / CLIP-ViT (LC) , CLIP4Clip , CLIPBERT , DALL·E 2 , DALL·E 3 , DeepSeek-VL2 , Eurus-RM-7B , Falcon , FLUX / FLUX1 , Gemini , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GPT-2 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-J , iFlytekSpark-13B , InstructBLIP , Internlm2-Reward , InternVideo , InternVL2 , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama-3-2-Vision , LLaVA , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Med , LLaVA-NeXT / LLaVA 1.6 , Med-Flamingo , Merlot Reserve , MiniGPT-4 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Molmo , mPLUG-2 , mPLUG-Owl3 , Nova Canvas , O1 / OpenAI-o1-preview , OpenCLIP , OPT , Otter , PaLM 62B , Pythia , Qwen , Qwen-VL , Qwen2.5 , RadFM , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , Stable Diffusion , Swin Transformer , Sybil , TerraMind , Tulu 2 , UniPerceiver , UniVL , Vicuna , Video-LLaMA , VideoCLIP , VindLU , VioLET , X-CLIP