Fuyu

text, image · generative · anchor

Vision-language model that feeds image patches straight into the language model without a separate vision encoder.

Note
the anchor is the release post the citing paper cites, not a paper: this model has none. The URL came out of the citation split across a line break and was rejoined; it answers 403 to an automated request, so it was not confirmed by fetching
Variants
Fuyu-8B

Findings

Shared mechanisms