CogVLM2

text, image · generative · anchor

Vision-language model on a Llama 3 backbone, cited in the corpus at 19B.

Note
anchor found by search and checked against this entry's own description: "CogVLM2: Visual Language Models for Image and Video Understanding"
Variants
CogVLM2-Llama3-19B, CogVLM2-19B

Findings

Shared mechanisms