Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a new embedding model designed to retrieve not just the answer to a query, but the surrounding context needed to verify it. The model introduces a training method that moves beyond the traditional "gold passage" approach, where systems typically identify only a single relevant chunk of text. By teaching embeddings to recognize supporting evidence, the update aims to reduce the ambiguity and token overhead faced by search agents and professionals who need to confirm facts within long documents.
The release coincides with context-bench, a new private benchmark created by turbopuffer to measure this specific capability. The benchmark includes 2,099 queries across 38,894 long documents in 21 domains, such as legal contracts, corporate filings, and clinical research.
### Beyond the single answer
Most retrieval systems break long documents into smaller chunks for easier indexing. This creates a problem: a chunk might contain the correct answer but lack the definition or background information located elsewhere in the document. When an AI agent retrieves only that isolated chunk, the result can be ambiguous or unverifiable without reading the entire source file.
Traditional training methods rely on "gold-chunk" annotations, where a large language model identifies the single best passage for a query. This approach treats every other part of the document as irrelevant, even if those parts provide necessary context. It also generates binary labels, which fail to capture the varying degrees of relevance between answer-bearing text and supporting evidence.
The new model addresses this by using a context compression model as a teacher. This teacher reads the query and document together, assigning a relevance score to every token. The embedding model then learns to aggregate these scores, allowing it to distinguish between chunks that contain the answer and chunks that support the answer, rather than discarding the latter as negatives.
### Technical approach and efficiency
The training process combines two loss functions: a document-level contrastive loss and a chunk-level distillation loss. The distillation loss teaches the model to match the teacher's continuous relevance scores. This means the model learns to retrieve support chunks with appropriate weight, rather than treating them as incorrect.
A key advantage of this method is its inference cost. The context compression model is used only during training. At runtime, the embedding model produces one vector per chunk with no additional latency or storage overhead. The released preview supports 1024-dimensional and int8 embeddings, making it compatible with existing vector databases.
The model was trained on roughly 430 public and in-house datasets covering over 50 languages. None of the training data included manual chunk-level annotations; all within-document supervision came from the automated teacher model. This reduces the cost and time associated with scaling training datasets, which typically requires an LLM to process each document individually.
### Evaluating context-aware retrieval
Standard benchmarks like ConTEB often assume a single relevant chunk per query, which limits their ability to measure context-aware retrieval. context-bench was designed to test three specific capabilities: document disambiguation, answer retrieval, and evidence recall.
Document disambiguation tests whether a model can identify the correct document among near-identical alternatives, such as lease agreements that differ only in tenant names or dates. Answer retrieval measures whether the model finds the chunk containing the direct answer. Evidence recall checks if the model also retrieves the supporting sentences necessary to verify that answer.
The benchmark uses sentence-level chunks to create a controlled environment. For each query, the primary target documents are intentionally long, with a median length of roughly 6,100 tokens. This ensures that the model must use information from distant parts of the document to succeed. The evaluation metrics, including Document@K and Evidence Recall@K, decompose the retrieval challenge into finding the right document, finding the right answer, and recovering the minimal context.
### Why this matters for software engineers and researchers
For developers building RAG (Retrieval-Augmented Generation) pipelines, this shift in embedding training offers a path to more reliable agents. Current systems often hallucinate or provide unverifiable answers because they retrieve isolated facts without their contextual anchors. By integrating models like pplx-embed-v2-context-9b-preview, engineers can reduce the token count sent to the final LLM while increasing the accuracy of the retrieved context.
Researchers working with long-form documents, such as scientific papers or legal contracts, benefit from the improved disambiguation capabilities. The model's ability to distinguish between near-duplicate documents helps prevent the retrieval of incorrect versions or outdated clauses. Access to the model's architecture and the public availability of the preview on Hugging Face allows these professionals to test the limits of context-aware retrieval in their own workflows. You can explore more about AI Engineering Courses to see how these tools fit into broader development practices.
Your membership also unlocks: