系列:Multimodal Decoding Notes

Multimodal Decoding Notes (2): VLM Evolution — ViT → CLIP → LLaVA

1. Why VLM Is the Pivot Chapter

Part (1) said: the essence of multimodality is alignment. The text side already has mature language models; what about the image side? The evolution of VLM is exactly the step-by-step process of making “image” digestible for a language model.

2. ViT: Slicing an Image into a Token Sequence (Dosovitskiy et al., 2020)

3. CLIP: Squeezing Image & Text into One Space via Contrastive Learning (Radford et al., 2021)

4. LLaVA: the Standard VLM Paradigm Is Set (NeurIPS 2023 Oral)

LLaVA proved one thing: you do not need to train a multimodal LLM from scratch — freeze the vision encoder + a light projection + fine-tune the LLM is enough.

LLaVA architecture: input image encoded by frozen CLIP ViT-L/14, bridged via a trainable projection to Vicuna's text-embedding space, concatenated with text tokens and autoregressively decoded by the LLM

Fig: The LLaVA standard paradigm — frozen vision encoder + light projection bridge + LLM fine-tuning (redrawn)

This paradigm (frozen vision encoder + projection bridge + LLM fine-tuning) spawned LLaVA-1.5 / NeXT, CogVLM, InternVL, Qwen-VL — almost every “chat about an image” model today descends from this skeleton.

5. What Determines a VLM’s Quality

6. Bridge to the Next Chapter

VLM solved “see + speak”. But many real tasks demand “act” — turning visual understanding into actions. Next we return to the other generative branch: how images / video are generated (diffusion and GAN); then we connect all this to robot actions, entering VLA and world models.

Practical tip: when looking at a multimodal model, ask three questions first — “whose vision encoder (CLIP/ViT family?), how was the alignment data made, how is the projection connected?” — these three quickly reveal engineering maturity, matching the evaluation framework from Part (1).

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。