Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Google DeepMind · Contributor · May 2026

Gemini Embedding 2 embeds video, audio, image, and text into a single shared representation space, so arbitrary interleaved combinations of those modalities can be compared directly. It reaches state-of-the-art results on unimodal, cross-modal, and multimodal retrieval benchmarks.