LLaVA-1.5 / LLaVA-v1.5

text, image · generative · anchor

Open-weight vision-language model pairing a CLIP vision encoder with an open language model through a projection layer, cited in the corpus at 7B and 13B.

Note
anchor found by search rather than in a citing paper, and checked against this entry's own description before it was recorded: "Improved Baselines with Visual Instruction Tuning"
Variants
LLaVA-1.5 7B, LLaVA-1.5 13B, LLaVA-v1.5-7B, LLaVA-v1.5-13B

Findings

Shared mechanisms