mPLUG-Owl3

text, image · generative · anchor

Vision-language model on a Qwen2 backbone, cited in the corpus for long image sequences.

Note
anchor found by search and checked against this entry's own description before it was recorded: "mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models"

Findings

Shared mechanisms