BLIP-2

text, image · generative · anchor

Vision-language model bridging a frozen image encoder and a frozen language model through a querying transformer.

Note
anchor found by search and checked against this entry: "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models", whose lightweight Querying Transformer is exactly what this entry describes
Variants
BLIP2-FlanT5-XL

Findings

Shared mechanisms