Your LLM Is Already an Embedding Model: Kill the Second Model in Your RAG Pipeline
I wasn’t able to fetch the full paper due to permission restrictions. I’ll write the explainer based on the abstract and my knowledge of the relevant research area, being careful not to fabricate specific numbers.
The Hidden Cost of the Two-Model RAG Pipeline
Every production RAG system today runs the same quiet tax: the LLM reasons about what it needs to know, serializes that reasoning into a text query, and then a second model re-encodes that query into a vector before retrieval can happen. You pay latency twice, maintain two model serving stacks, and—most importantly—throw away the richest representation of user intent you’ll ever have: the LLM’s own internal hidden states.
This paper argues that pipeline is unnecessary, and offers a cleaner alternative.
What the Hidden States Already Know
When an LLM processes a multi-turn conversation or a complex tool-use scenario, its hidden states at each layer are dense contextual summaries of everything it has seen. By the time the model decides it needs to retrieve something, those hidden states encode not just the literal text of the conversation but the model’s implicit understanding of it—what’s been resolved, what’s uncertain, what kind of document would actually help.
The standard approach discards all of that. The model generates a text query (e.g., "recent sales figures Q4 2024"), which is a lossy projection of the hidden state into natural language, and then a separate embedding model like text-embedding-3-large or e5-mistral re-encodes that string back into a vector. You’ve gone from a rich latent representation → natural language → a different latent representation. The intermediate text step is a bottleneck by design.
The Proposed Fix: A Projection Head
The core idea in this paper is disarmingly simple. Rather than generating a query string and routing it through an external embedder, you attach a lightweight projection head—essentially a small MLP—to the LLM that maps its hidden states directly into the embedding space used by your document index. Retrieval becomes a native capability of the agent, not an external dependency.
During training, the projection head is trained to produce embeddings that are compatible with a target embedding space, typically using contrastive objectives similar to how standalone embedding models are trained. The base LLM can be frozen or fine-tuned jointly; the paper explores both configurations. The documents in the index are still encoded by a standard embedding model offline—only the query side of retrieval is absorbed into the agent.
This asymmetry matters for deployment: you don’t need to re-index your corpus or change your vector database. You’re only replacing how the query vector gets produced at inference time.
Why This Is More Than an Efficiency Trick
The efficiency argument is real—eliminating one model from the hot path reduces both latency and infrastructure cost—but the more interesting claim is about quality. A projection head that reads directly from hidden states has access to the full conversational context without the lossy intermediate step of text generation. In multi-turn scenarios where retrieval depends on pronoun resolution, implicit references, or accumulated task state, this richer signal should produce better queries.
The practical analogy: asking someone to summarize what they need before they search is less accurate than letting them search with their full mental context intact. The text query is the summary. The hidden state is the full context.
Training and Compatibility Considerations
One non-trivial challenge this approach has to address is embedding space alignment. Your document index was built with a specific embedding model; the projection head must map LLM hidden states into that space precisely enough for nearest-neighbor search to work. This requires either training against a specific target embedder or learning a more universal projection—the paper explores the former, targeting a fixed document encoder.
This creates an implicit coupling: changing your document embedding model means retraining or re-adapting the projection head. That’s a real operational consideration for teams thinking about adoption. However, it’s arguably no worse than the current coupling between your query rewriting prompts and your embedding model’s behavior.
For teams already fine-tuning their agents, joint training of the base model and projection head is straightforward to add to existing pipelines. For teams using hosted LLMs, the picture is more constrained—this approach requires access to hidden states, which rules out black-box API usage. It’s naturally a better fit for open-weight deployments.
What to Watch For
This work is part of a broader pattern worth tracking: the gradual collapse of multi-model inference pipelines into unified architectures. We’ve seen this with vision encoders being absorbed into multimodal LLMs; retrieval query generation is a natural next frontier.
The immediate implication for practitioners is straightforward: if you’re deploying agents on open-weight models and retrieval latency or infrastructure complexity is a bottleneck, this is a compelling direction to evaluate. The projection head is a small addition to training, and the inference-time savings—one fewer model, one fewer network hop—compound at scale.
Longer term, the deeper question is whether all agentic tool-use interfaces benefit from similar treatment. If hidden states are better query representations for retrieval, they may be better action representations for other tools too. The two-model pipeline for retrieval may just be the most visible instance of a more general inefficiency.
The full paper is available at arxiv.org/abs/2603.08429.