Light Dark Meta · 2024-07 · text · generative · anchor
Open-weight language model family extending Llama 3 with a longer context window, cited in the corpus in its 8B and 70B sizes and in the instruction-tuned 8B checkpoint.
Variants Llama 3.1 8B , Llama 3.1 8B Instruct , Llama 3.1 70B , Llama 3.1 405B , Llama 3.1 70B Instruct , Llama 3.1 405B Instruct , Llama 3.1 Instruct , Llama 3.1-3B , Meta-Llama-3.1-70B-Instruct , Llama-3.1-Nemotron-70B-Reward , Meta-Llama-3.1-8B-Instruct Findings IC-007 Most LLMs do not align closely with human moral preferences on multilingual trolley problems IC-008 Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances IC-009 LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages IC-011 Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. IC-013 Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models. IC-017 Truncating MLP weights improves few-shot Chain-of-Thought reasoning accuracy on GSM8K for Phi-3 and Llama-3.1-8B IC-029 Large language models show conformity to group answers in multi-agent interactions IC-030 Larger language models exhibit higher independence rates and lower conformity under some protocols IC-072 LLM agents of varying scales exhibit a failure mode on web automation tasks when processing raw, complex web page observations, with the penalty being more severe for smaller models IC-073 Released LLMs (GPT-4o, Llama-3.1-70B, Qwen2-7B, etc.) show limited workflow orchestration capability that degrades as workflow complexity increases IC-074 Released LLMs achieve F1 plan scores between 42.7 and 86.7 on the T-Eval plan task IC-082 GPT-3.5, GPT-4o, Claude-3.5-Sonnet, and Llama-3.1-8B produce explanations on the BBQ social bias task that are systematically unfaithful for identity and behavior concepts while remaining faithful for context concepts, with specific patterns of hiding safety-measure influence and social bias IC-087 Answer symbol production in OLMo 7B Instruct, Llama 3.1 8B Instruct, and Qwen 2.5 1.5B Instruct is causally attributed to a few middle layers and specifically their multi-head self-attention mechanisms, with a sparse set of 1-4 attention heads per layer responsible IC-092 All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cue IC-093 No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent events IC-094 Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMs IC-095 Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering IC-096 Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top IC-097 Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B IC-103 LLMs' value rankings align with the universal human value hierarchy under most prompting conditions IC-104 Value anchor prompting produces LLM value correlation structures that closely match the human circular value structure, while standard prompting does not IC-105 Value anchoring produces a sinusoidal scoring pattern around the circular value structure, with scores decreasing as circular distance from the anchor increases IC-112 Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increases IC-113 Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputs IC-114 Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl range IC-166 A 1-dimensional subspace in a single layer encodes the context-versus-prior decision in Llama-3.1-8B, Gemma-2 9B, and Mistral-v0.3 7B, and setting this subspace steers the released (non-fine-tuned) models' behavior IC-175 Context-sufficiency performance is scale-dependent: larger LLMs achieve high accuracy with sufficient context but still answer correctly 35-62% of the time without it, while smaller models hallucinate or abstain even with sufficient context IC-176 LLaMA 3.1 8B Instruct's KGQA accuracy degrades with increasing numbers of retrieved triples, while GPT-4o-mini's accuracy improves, revealing different context-handling capacities IC-188 LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categories IC-189 Most LLMs show degraded instruction satisfaction when user instructions conflict with system messages, indicating difficulty in prioritizing system message constraints IC-190 LLMs show progressive degradation in system message constraint following across multi-turn conversations, with dependent conversations degrading faster than parallel ones IC-191 Attention allocated to system messages correlates with following ability, and models do not strictly distinguish system from user messages based on marker tokens IC-202 All six evaluated LLMs achieve very low accuracy on OpenRCA, with no model solving any three-element root cause query IC-203 Gemini 1.5 Pro's RCA-Agent accuracy drops 68.4% when code execution fails, far exceeding the drops for Claude 3.5 (17.9%) and GPT-4o (15.6%) IC-205 Gemma 2's SAE features exhibit depth-dependent organization, with polysemantic features in early layers and persistent, matchable features in later layers IC-221 GPT-4o ReAct success rate drops from 47% on synchronous to 11% on asynchronous planning tasks, and all other tested LLMs show equal or worse performance IC-242 Most LLMs exhibit higher bias ratios in multi-turn dialogues than in single-turn, with bias accumulating across successive turns IC-244 No LLM demonstrates consistently strong fairness across both comprehension-focused and bias-resistance multi-turn tasks; models show complementary failure patterns IC-255 GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlations IC-264 All 18 evaluated LLMs fail to abstain when the provided context lacks the answer, with performance gaps of 13.6% to 68.4% relative to the original context IC-265 Model families show extreme variation in detecting conflicting answers in inconsistent contexts, with phi-3 series at 5.8% average accuracy versus GPT-4 series at 89.35% IC-266 GPT-4o drops from 96.3% closed-book accuracy to 47.5% when given counterfactual context that contradicts its parametric knowledge, far below the 95% human accuracy on the same items IC-294 Llama-3.1-8B-Instruct with 2-shot prompting achieves limited rationale extraction quality (F1 15.7–48.3) across four text classification datasets IC-298 Weight similarity in open-source LLMs is organized in a depth-dependent structure with adjacent-layer similarity and distinct clusters at specific depths IC-299 Instruction tuning preserves the weight-matrix structure of LLMs, with DOCS scores exceeding 0.7 across all matrices IC-302 Llama-3.1 models perform between unigram-inference and bigram-inference on Markov chain ICL, with performance improving monotonically with model scale IC-303 Explicitly stating the Markovian structure in the prompt significantly improves Llama-3.1-70B's next-state prediction on the Markov chain task IC-315 CoT prompting (reasoning + instruction) yields larger relative gains for larger LLMs and harder problems in competitive code generation, with the effect reversing for the most capable models IC-316 Multi-turn code generation without CoT degrades performance for smaller Llama models and GPT-4o compared to single-turn repeated sampling under equal compute budgets IC-317 More detailed execution feedback (LDB) induces exploitative behavior in Llama 3.1 models, reducing code diversity and hurting performance at large sample budgets IC-340 GCG jailbreaking attacks exhibit strong model-specific transferability, achieving below 3% ASR on Llama-2-13b-chat and Llama-3.1-8b-instruct but above 90% ASR on Vicuna-13b-v1.5 and Mistral-7b-instruct IC-351 GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarks IC-352 RAG with sufficient retrieved tokens outperforms direct long-context for Qwen2-72b-instruct on >100k tasks, while at 32k the default RAG setting underperforms direct long-context for GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct IC-353 Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval test IC-393 GPT-4-turbo, Llama-3.1-8B-Instruct, and OpenAI Moderation show declining hate speech detection accuracy as sentence implicitness increases, with very low success rates in the highest implicitness ranges IC-421 Sequential context-switching queries jailbreak Llama and Mistral models at 95% attack success rate IC-422 Safety fine-tuning in Llama models improves with parameter size but exhibits diminishing returns IC-423 Llama-3.1-8B-Instruct achieves 97.8% accuracy as a zero-shot toxicity classifier on ToxiGen, outperforming Llama-3-Guard-1B and matching Llama-3-Guard-8B IC-430 Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal IC-433 LLMs exhibit a non-monotonic ID-OOD performance gap (generalization valley) that peaks at intermediate task complexity IC-434 The critical complexity at which LLMs over-rely on memorization shifts to higher task difficulty as model size increases IC-467 Llama-3.1-405B's standard speculative decoding verification rejects correct continuations from GPT-4o, Llama-3.1-8B, and human text, accepting only roughly two tokens before the first rejection for GPT-4o IC-468 Llama-3.1-405B's last hidden layer embeddings of erroneous tokens contain a linearly detectable error signal that a simple logistic regression head can exploit to flag incorrect continuations IC-474 GPT-4, GPT-4o, and Llama-3.1-405B fail at knowledge classification and comparison without chain-of-thought IC-475 GPT-4, GPT-3.5, GPT-4o, and Llama-3.1-405B fail at inverse knowledge search regardless of prompting IC-477 Released LLMs achieve limited success rates as web agents on WebArena-Lite, with open-source models substantially below proprietary ones IC-481 Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increases IC-482 In Llama-3.1-8B, in-context graph structure is absent in early layers dominated by semantic priors and emerges clearly in deeper layers IC-483 When in-context graph structure conflicts with pretrained semantic priors, Llama-3.1-8B encodes the in-context structure in higher principal components while the semantic prior dominates the first two IC-484 Re-scaling the first 2-3 principal components of Llama-3.1-8B token representations causally shifts next-token predictions toward the target graph position IC-500 GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B solve multi-step retrieval problems without fine-tuning, achieving near-perfect accuracy for chains of up to 5 steps IC-539 GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding types IC-540 Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models IC-541 Code-based output (Python or TikZ) produces images with notably higher scientific style scores than direct image generation from Stable Diffusion and DALL-E IC-547 Llama3.2 3B and Llama3.1 8B exhibit saturation events in which the top prediction, once it appears at a given layer, remains unchanged through all subsequent layers IC-549 All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench IC-550 Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B models IC-552 GPT-4, Llama-3.1-8B, and Qwen-2-72B all improve on ALFWorld and WebShop when given a generated workflow as structured prior knowledge IC-575 Four released LLMs (LLaMA-3.1-8B, Mistral-7B, Qwen2-7B, Yi-1.5-9B) can perform in-context learning on continuous vector representations projected into their embedding space, matching or outperforming few-shot ICL across text, time-series, graph, and fMRI tasks IC-576 For 10-digit numerical function regression, vector-ICL consistently outperforms few-shot ICL with raw number inputs across all four LLMs because continuous representations avoid multi-token splitting Shared mechanisms Circular representation also in Gemini 1.0 Pro , Gemma 2 , GPT-2 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , TerraMind Depth-dependent structure also in Baichuan 2 , BERT , BLIP-2 , BLOOM , Chameleon , CLIP / CLIP-ViT (LC) , DeepFloyd IF , DeiT-III , DINO , DINOv2 , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-J , GPT-NeoX-20B , Griffin , I3D , Idefics , InstructBLIP , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.2 , Llama-3.2-3B , LLaVA , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Phi , MAE , MAE-B/16 , Mamba , MiniGPT-4 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , mPLUG-Owl , MPT , MultiBERTs , MViT V2 , OLMo / OLMo base , OpenCLIP , OPT , Phi-2 , Pythia , Qwen 2 , Qwen2-VL , Qwen2.5 , RoBERTa / RoBERTa-L , RWKV , SALMONN , SAM , SlowFast , Stable Diffusion , Swin Transformer , TerraMind , TimesFormer , TSM , Uniformer , Vicuna , VideoMAE , ViT , X3D , Yi Distance preservation also in CLIP / CLIP-ViT (LC) , CoPlace , DINO , DINOv2 , Gemma , Gemma 2 , ImageBind , LanguageBind , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Med , MAE , MAE-B/16 , OpenCLIP , Pythia , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SigLIP , SLIP , TerraMind , ViT Explanation faithfulness also in BakLLaVA , CF2 , Claude 3 , Claude 3.5 , CLIP / CLIP-ViT (LC) , DRUM , Fuyu , GEM , Gemini 1.5 / Gemini Pro 1.5 , Gemma 2 , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-J , Idefics , LLaVA-NeXT / LLaVA 1.6 , MAE-B/16 , MobileNetV2 , mPLUG-Owl3 , OpenFlamingo , PGExplainer , Pythia , RCExplainer , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SigLIP , SigLIP-2 , Stable Diffusion , TAGExplainer , ViT Failure mode also in AASIST , ADM , Aegis-Guard-Defensive , Alpaca , AnyLoc , AutoTikZ / DataTikZ , Baichuan , Baichuan 2 , Baichuan2-13B , BakLLaVA , Bard , BEiT , BERT , BingChat , BLIP , BLIP-2 , BLOOM , BSRNet , CF2 , Chat-UniVi-7B , ChatGLM-6B / ChatGLM-6b-2 , ChatGLM2 , ChatGPT , CLAP , Claude 1.3 , Claude 2.0 , Claude 2.1 , Claude 3 , Claude 3.5 , CLEAR , CLIP / CLIP-ViT (LC) , CLIP4Clip , CLIPBERT , CLIPCap , CLMBR-T-BASE , CloFNet , Code Llama , CodeGeex2 , CodeGen , CodeLlama-13B , CodeLlama-34B , CogVLM2 , Cohere Command R , CoMEt , Command R+ , CONCH , CycleGAN , DALL-E , DALL·E 2 , DALL·E 3 , DASR , DECAF , DeepSeek-2-Chat , DeepSeek-2-Coder , DeepSeek-V2-0628 , DeepSeek-VL , DeepSeek-VL2 , DeiT , DeiT-III , Depth Anything , DETR-R50 , DimeNet++ , DINO , DINO-FocalNet-Large , DINOv2 , EGNN , Emu2 , EquiformerV2 , ESCN , ESM-2 , ESM3 , ESRGAN , EVA-CLIP , EVE , Falcon , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 , FLAN-T5 , Florence-2 , FLUX / FLUX1 , FoveaBox , Fuyu , Galactica-6.7B , GAT , 3D Gaussian Splatting , GCN , Gemini , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GIN , GLIDE , GLM-4 , GLM-4V , GloVe , GP-UNIT , GPT-2 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4.1 , GPT-4o , GPT-J , GPT-NeoX-20B , GraphSAGE , Grounding DINO , Guanaco , GVP , Hawkeye , HiFaceGAN , HPS , HuggingGPT , IDDPM , Idefics , Idefics2 , ImageBind , ImageBind-LLM-7B , Imagen Video , ImageReward , 12-in-1 , InstructBLIP , InstructGPT , InternLM-2.5-7B , InternLM-XComposer2-VL , InternVideo , InternVL-1.5 , InternVL2 , Koala , LegalBERT , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.2 , Llama-3.2-3B , Llama-3-2-Vision , LLaMA-Adapter v2 , Llama Guard , Llama Guard 2 , Llama-Guard 3 , Llama-VID , LLaVA , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Med , LLaVA-NeXT / LLaVA 1.6 , LLaVA-OneVision , LongVA-7B , LOVT , LWM-1M-JAX , MACE , MAE , Med-Flamingo , Merlot Reserve , MGCA , MiDaS , MiniCPM-V , MiniGPT-4 , Mip-Splatting , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mistral Large 2 , Mistral Large V2 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , MobileNetV2 , Molmo , MolmoE-7B , Momentor , Moondream2 , Moonshot-v1-8k , mPLUG-2 , mPLUG-Owl , mPLUG-Owl3 , mPLUG-Owl2 , MPT , MSA Transformer , MultiBERTs , Nova Canvas , Nova Lite , Nova Pro , O1 / OpenAI-o1-preview , O3 , O4-mini , OLMo / OLMo base , OneLLM , OpenAI Moderation , OpenChat-3.5-0106 , OpenCLIP , OpenFlamingo , OPT , Otter , Otter-7B , PaLM 2 , PaLM 62B , PandaGPT-7B , PerSAM , Phi-3 , Phi-3.5 Mini Instruct , PickScore , PLIP , Prismatic , ProGen-2 , Pythia , Qwen1.5 , Qwen 2 , Qwen 2.5 72B Instruct , Qwen2-VL , Qwen-Audio , Qwen-VL , Qwen2.5 , Qwen2-Audio , R2D2 , RadFM , RCExplainer , RD-SR , RealESRNet , Reprover , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , RetinaNet-R50 , RivaGAN , RS-LDS , RT-1 , RT-2 , SALMONN , SAM , SAM 2 , SAULLM 54B , Scaffold-GS , SchNet , Seed-LLaMA-8B , SGC , SIREN , Sketch Transformer , SLD-max , SLD-medium , SLD-strong , SLDS , SLIMG , SpeechGPT , SphereNet , SRResNet , Stable Diffusion , StackLLaMA , Starcoder , StegaStamp , StyleGAN2-ADA , Swin Transformer , T5 , TD-MPC , TerraMind , TimeChat , TranceptionEVE , TreeRing , Tulu 2 , UnifiedQA , UniPerceiver , UNITER , UniVL , Van Gogh Diffusion , VERA , VGG / VGG13 , Vicuna , Video-Chat-7B , Video-ChatGPT , Video-LLaMA , Video-LLaMA-2-13B , Video-LLaVA , VideoCLIP , ViLA-8B , ViLBERT , VindLU , VioLET , ViRTex , ViT , ViV1T , VTG-LLM , WildGuard , Wizardlm , X-CLIP , X-InstructBLIP-7B , XGen-MM , Xlm-R , Zephyr-7B-beta Linear representation also in BLOOM , Cambrian-1 , Chameleon , CLIP / CLIP-ViT (LC) , DINOv2 , EVA-CLIP , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-J , HPSv2 , ImageBind , InstructBLIP , LanguageBind , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 , MAE , Mamba , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , OLMo / OLMo base , OpenCLIP , Phi-3 , PickScore , Pythia , Qwen2-VL , Qwen2.5 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SALMONN , SAM , SigLIP , TerraMind , Tulu 2 , Vicuna , ViT Method artefact also in Baichuan , CLIP / CLIP-ViT (LC) , ConvNeXt , EfficientNet , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-3.5 / ChatGPT-3.5 , GPT-4o , GPT-J , InternLM , Llama 2 / Llama 2 base , Llama 3 , MAP-NEO , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , OLMo / OLMo base , OpenCLIP , OpenLLaMA , OPT , Pythia , Qwen1.5 , Qwen 2 , Qwen2.5 , RedPajama , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , Skywork , Stable Diffusion , StableLM , TerraMind , ViT , Yi , ZiYA2 Scale-dependent behaviour also in Aquila-7B , BEiT , BERT , BLOOM , Claude 2.1 , Claude 3 , Claude 3.5 , CLIP / CLIP-ViT (LC) , Code Llama , CodeGen , Cohere Command R , DeepSeek LLM , DeepSeekMoE , DeiT-III , DINO , DINOv2 , EquiformerV2 , ESCN , Falcon , FLAN-T5 , Gemini 1.0 Pro , Gemini 1.5 / Gemini Pro 1.5 , Gemma , Gemma 2 , GPT-2 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-J , GPT-Neo , I3D , Idefics , InternLM-2.5-7B , InternLM-XComposer2-VL , InternLM2 , InternVL-1.5 , InternVL2 , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.2 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 , LongVA-7B , MAE , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Moirai , MPT , MViT V2 , O1 / OpenAI-o1-preview , OLMo / OLMo base , OpenCLIP , OpenFlamingo , OpenLLaMA , OPT , PaLM 2 , Phi-3 , Platypus2-Instruct-70B , Pythia , Qwen , Qwen1.5 , Qwen 2 , Qwen 2.5 72B Instruct , Qwen2-VL , Qwen-Audio , Qwen2.5 , RedPajama-INCITE , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SlowFast , Solar 10.7B , Stable Diffusion , StableLM , Swin Transformer , T5 , TerraMind , text-ada-001 , TigerBot-70B-Chat , TimesFormer , TSM , Tulu 2 , Uniformer , Vicuna , VideoMAE , ViLA-8B , Wizardlm , X3D , XGLM , Xwin-LM-70B-v0.1 , Yi