Retrieval Semantics: HyDE vs Query Rewriting
Any RAG system is built to prevent this by providing additional context beyond just what the model has learned in its training dataset. Doing it means retrieving this context reliably, including it in the prompt, and passing it to the model. As simple as that sounds, it introduces another problem layer that can condemn any RAG system to an execution by hanging (plus the creative profanity that anyone who's ever been stuck in that system will most likely utter).
Victor Ohachor ·

But the focus of this writing is not to explain what RAG means. There are a number of articles that do a better job at this. A non-exhaustive list:
- IBM on Retrieval Augmented Generation
- AWS on Retrieval Augmented Generation
- Prompt Engineering Guide wrote an illustrative article called RAG for LLMs
In fact, a simple Google search for the word RAG throws a list of high quality articles at you. The recent development in AI since OpenAI first launched ChatGPT has taken over the original meaning of the keyword.
Yet this article looks deeper into RAG techniques, particularly at the retrieval layer. It assumes that you are comfortable with the word "RAG". Comfortable in the sense that you don't need any mansplaining from me to understand that retrieval is the backbone of any RAG system.
If you checked this simple box, follow me as we journey through a core problem engineers building RAG systems face in production.
What is this "Core Problem?"
That retrieval without a testimony to semantic awareness invalidates the benefits RAG presents in any AI system.
LLMs generate content; the kind of content depending on the kind of model. This is a game-changer if you can assert that generation is deterministic. But it isn't. And this is the problem that any RAG system targets, yet not the core problem I inferred above.
LLMs hallucinate; and prompt engineering alone is powerless against this problem. Say a LLM generates text, it achieved it by predicting next token, given previous tokens. And these tokens are already available in its corpus. How common a combination of tokens is depends on how many times it appears in the training data. With an instruction-tuned LLM, any user prompt that requests for information that is not in the model's training data causes a glitch that makes the model to make up facts.
Any RAG system is built to prevent this by providing additional context beyond just what the model has learned in its training dataset. Doing it means retrieving this context reliably, including it in the prompt, and passing it to the model. As simple as that sounds, it introduces another problem layer that can condemn any RAG system to an execution by hanging (plus the creative profanity that anyone who's ever been stuck in that system will most likely utter).
Usually retrieval is done by searching semantically through a knowledge base the RAG system tracks for something that directly relates to the user's prompt or question.
Searching semantically means that it meaning behind the user's prompt is what is checked against the knowledge base by the RAG system. The RAG system, during the retrieval stage of the pipeline, solely focuses on retrieving documents that are relevant to the user's question or prompt based on the connotation with what the knowledge base already supports.
It is important you understand the last paragraph because that leads us directly to the core problem:
That begs the question...
Is Semantic Awareness Situational Awareness?
In social context, situational awareness is the perception of one's environment and the ability to make informed predictions based on that perception (Mohn, 2019). It involves scanning your environment and documenting your observations, comprehending the meaning behind it, and making informed decisions based on your understanding.
Interestingly, as layered as these steps look, it can happen on the fly. In fact, it is supposed to happen on the fly. For example, when conversing with a friend in real life, you observe their body language, the tone of their voice, their facial expressions, comprehend the meaning, and align your responses based on that.
Compare that to semantic awareness for a moment...
The ability to interpret meaning in context rather than just matching keywords. In NLP, this is often called semantic understanding.
If you isolate any piece of text, say a user's question, unless they provide additional information, you will find that what they inherently mean changing based on the context. A more practical example:
A user asks a model, "How are you doing?"
If it is the first message in that chat, the model will most likely respond that it is doing well. Gemini for one says, "I'm doing well, thanks for asking! What's on your mind today?"
But if the previous conversation included the user saying, "Whenever I ask you how you are doing, reply with how the world is doing." Now Gemini responds with:
The world right now is navigating a fascinating transitional era. We are seeing rapid, sweeping technological breakthroughs—especially with deep AI integration across daily life and major strides in renewable energy infrastructure—that are genuinely pushing the boundaries of medicine and global transit. At the same time, the global community continues to grapple with the complex realities of climate adaptation, economic shifts, and finding equitable, sustainable ways to manage this breakneck pace of progress. It is a complex, highly dynamic picture: plenty of hurdles to overcome, but an incredible amount of human ingenuity hard at work.
Funny, right?
By default, LLMs have semantic awareness, but only when you have presented them with the observations. They cannot observe by themselves if information about a user's question is not in their training data. They need a layer that gathers the context they need to make decisions on a user's question.
That is semantic awareness, where the context is already available. Unless situational awareness where the context (also called observations) is first done. There is the answer!
Hypothetical Document Embeddings (HyDE)
A fact that has been established in this article is that for a RAG system to be effective, it has to retrieve the right context. If the context is wrong, the response will be irrelevant to the users. Maybe annoying.
Also remember that retrieval works based on semantics.
But a question and an answer are in different worlds and often do not look alike especially when the question is short. Say a child comes up to you and ask, "Why is he obese?" Your mind will likely stream through a number of new questions. What your brain is doing is searching through your memory to connect the dots between the core keywords in the question.
Questions like "Who is 'he'?", "What does 'obese' mean?", "Why is he asking?", and more. You are trying to gather enough context to answer the child's question.
For a RAG system, the retrieval layer often uses semantic search. First, there is a knowledge base possessing proprietary or domain-specific knowledge that is useful to the users of that system. Then, there is a search layer where the system embeds the user's question into vectors and performs a similarity search that checks for related documents. These documents have already been embedded too.
This is why we stated somewhere in this article that the retrieval layer is not trying to answer user's question, it only looks up relevant documents based on the user's question. With HyDE, this changes fast.
The system passes the user's question to a LLM and asks it to write an answer. Even if the LLM hallucinates or gets it wrong, the vocabulary and structural patterns of this "fake" answer will match the target document. Hence, the retrieval layer will most likely match the right context when it compares this fake answer against the knowledge base.
Note that the word, "document," anywhere in this article simply means a piece of text.
The core idea is simple: An answer looks like an answer. Hence, to find a true answer, we must search with what looks like the answer (Shekhar, 2026).
Where does HyDE Fail Us?
Think about it for a minute.
1...2...3...4...5...6...7...8...9...10...
Yeah, that one... Write it down.
11...12...13...14...15...16...17...18...19...20...
Exactly! Think deeply now about the question. Create a scenario to support your hypothesis and play out the consequences of some flaws you spotted.
Did you say that it was an expensive technique? Yeah, it is, but that is not it. Both retrieval techniques this article detailed are not cheap. Also, HyDE introduces latency because of the extra step of synthesizing a "fake" answer, but still not the core issue.
When you synthesize an answer before the true context is retrieved, you are essentially handing the steering wheel of your RAG system to the model's ungrounded imagination. This is known Premature Hallucination Amplication.
That is the core issue!
If a LLM had to based its "fake" answer on its training data, it could generate the perfect illusion that can cause the retrieval layer to drift from relevant documents, resulting in a context with inaccuracy that is worse than when you didn't apply HyDE technique. This is because the RAG's knowledge base is gullible. It doesn't understand the concept of truth. It understands geometric proximity.
Query Rewriting, A Sibling of HyDE
Rewrite your query; give it context! Don't leave it to chances that your RAG system do not fail.
Now that we have dissected HyDE and pulled out its guts, you are ready to face Query Rewriting.
Query rewriting means exactly what it sounds like. It is a technique that involves rewriting the user's question to give it deeper context. Unlike HyDE that can be also leveraged where you have a bare user's question that does not look like the the documents that answer it in the vector space, query rewriting takes a different approach.
Instead of synthesizing an answer before the real one, the system rewrites the query to include more context. This is particularly useful mid-conversations, where the question references prior context implicitly. Say the user asks, "what about pricing for that?." "That" has no vector space meaning of its own and requires that the system peeks into the previous conversations to establish the full form of the question.
There are three (3) concrete techniques of query rewriting, each solving a different flavor of the problem...
- Conversational condensation
- RAG-Fusion
- Step-back prompting
With conversational condensation, the system is simply combining the chat history and the new question to create a more robust question where the pronouns and references have been resolved.
RAG-Fusion will require a dedicated article. Hence, look forward to that article. But to simply put, instead of rewriting to one query, the system generates multiple variations of the user question, runs each against the retrieval layer, and combines the results using RRF (more on this in the next article). Example could be a prompt like "generate 3-5 different phrasings/angles on this question."
Step-back prompting is the process of raising more questions to ensure the system has enough context before the retrieval layer is concluded. Before answering a specific question, it first asks (and retrieves for) a more abstract, general version of it. Say the user asks, "What was the exact revenue of Company X in Q3 2019?", a step-back question could be "What is Company X's financial history?". The goal is to retrieve broad context from the step-back question first, then use that as grounding context for answering the specific original question.
If you need to revisit this article again to properly digest and retain this, please do. See you in the next article.
References
Mohn, E. (2019). Situational awareness. In Research Starters. EBSCO. https://www.ebsco.com/research-starters/social-sciences-and-humanities/situational-awareness/
Amit, S. (2026, July 6). How does HyDE work in RAG? https://outcomeschool.com/blog/how-does-hyde-work