Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ P
Is it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. It is not yet on the Perplexity API. The model card notes that weights and interface may change without backward compatibility.
Why the gold passage falls short
RAG systems split long documents into chunks. A chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address this with late chunking. The document is encoded in one pass, then pooled per chunk.
Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists 3 more problems. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. Labels are also tied to one chunking strategy.
How the training works
The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.
Chunk relevance: the mean of the top n token scores inside each chunk.
Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
Distillation loss: forward KL divergence between teacher and student distributions.
Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.
Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.
The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.
Interactive explainer
pplx-embed-v2 Contextual Embedding Explainer
ctx
How pplx-embed-v2-context retrieves answers and their evidence
An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.
01Why context
02How it trains
03Reported results
04Storage math
Three lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.
QUERYWhen does 5 Park Avenue’s lease end and what is the current rent?
Isolated sentencesContextual (late chunking)
Run retrieval
Pick a mode and press Run retrieval.
Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.
A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.
Chunk boundaries:
Sentences2-sentence chunks
Temperature0.25
Animate
1 Teacher scores tokens2 Top-n mean per chunk3 Softmax target4 Student matches via KL
Gold-chunk label (one-hot)
Teacher target (soft, from token scores)
Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.
context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.
pplx-embed-v2-context-9b-previewvoyage-context-4 (derived from reported gap)
Replay
Voyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.
Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.
Chunks
100,000
1,000,000
2,458,072 (context-bench)
10,000,000
100,000,000
1024-d2048-d
int8float32
0
vector storage (vectors only, not full index)
Bytes per vector
–
vs 2048-d float32
–
Chunk-size sensitivity (64 to 512 tokens)
81.0% to 79.9%
Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.
Sources: Perplexity Research · Model card
Built by Marktechpost