Small vision-language model in the LLaVA line built on a Phi language backbone, cited in the corpus in two generations.
Note
anchor found by search and checked against this entry: "LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model", which builds LLaVA-1.5 on a Phi backbone, exactly as this entry says