Perplexity releases multimodal embedding models that keep token-level detail
The MIT-licensed 0.6B and 9B retrievers embed text and page images into a shared multi-vector space, enabling a small query encoder to search a higher-quality 9B-built index.
2 min read
Perplexity released two MIT-licensed multimodal retrieval models, pplx-embed-v2-late-0.6b and pplx-embed-v2-late-9b, on October 7. Both are downloadable from Hugging Face and work with Sentence Transformers, but neither was available through a Hugging Face inference provider at launch.
Unlike dense retrievers that compress an entire query or document into one vector, the new models retain a 128-dimensional vector for each token and score query–document matches with MaxSim. This ColBERT-style late interaction preserves more local detail, at the cost of larger indexes and more retrieval computation than a single-vector system.
The two sizes share an embedding space. A team can therefore index documents with the 9B model and encode live queries with the 0.6B model, aiming to keep higher-quality document representations while lowering query-time cost. The models handle text and rendered page images, making them relevant to visual PDF search without OCR. They require separate text-only and image-only encoding calls; mixed text-plus-image batches are not supported.
Perplexity reports 62.3% nDCG@10 for the 0.6B model and 65.2% for the 9B model on the public image portion of ViDoRe V3. On the Markdown representation, it reports 61.2% and 64.7%, respectively. Those figures come from Perplexity’s own release evaluation rather than an independent benchmark submission. The company also notes that two Q2D-Web judgment sets are derived from documents surfaced by its own systems, which may favor similar ranking behavior; it emphasizes a broader combined judgment set to reduce that bias.
The model cards report 340 million active parameters for the smaller checkpoint and 7.4 billion for the larger one. Both were distilled from an internal 18B ColBERT teacher.
Sources: Perplexity release post · 0.6B model card · 9B model card · Unsplash image license
Conversation
0 approved commentsNo comments yet. Start a thoughtful conversation.
Join the conversation.
Sign in with ChatGPT to leave a comment or rate this article.