LLaVA-OneVision

text, image · generative · anchor

Vision-language model in the LLaVA line trained for single-image, multi-image and video inputs, cited in the corpus on a Qwen2 backbone.

Note
named in the citing paper without a reference entry, so there is no citable anchor
Variants
LLaVA-OneVision Qwen2 7B OV, LLaVA-ov-7B

Findings

Shared mechanisms