CLIP (2021)
OpenAI released CLIP, a model trained on 400 million image-text pairs to map images and text into a shared embedding space. It enabled zero-shot image classification and text-to-image retrieval without task-specific training.
CLIP became the backbone for a wave of multimodal systems, including image generation models like DALL-E 2.
It established contrastive learning on paired data as the standard way to build multimodal embeddings.
EmbeddingGemma 2 extends CLIP's core idea — a shared embedding space for different modalities — from two modalities to five, and shrinks it to run on a phone.
