Does your chatbot search for a direct fact, a comparison and a personal recommendation in exactly the same way? An adaptive RAG architecture classifies the question first, then chooses a search strategy that fits what the customer actually needs.
Basic RAG can answer direct questions well. It becomes less reliable when a customer asks the bot to compare plans, explain a choice or consider their particular situation.
The problem is not always the model. Sometimes the search method is wrong before the model even begins writing.
A basic RAG chatbot receives a question, searches for document chunks with similar meaning and sends the closest chunks to a language model.
That process works when the question has a direct answer in one place. It struggles when the answer depends on several documents, different viewpoints or context from an earlier message.
Example:
An online coaching platform loads its fee structure, course syllabus and refund rules into a basic RAG chatbot. A student asks, “What is the fee for the weekend batch?” and receives the correct amount from the fee document.
A parent then asks, “Is the weekend batch better than the regular batch for a working student?” The chatbot returns the same fee paragraph because it searches both questions through the same retrieval path.
The first question needs a precise lookup. The second needs a comparison across several pieces of content.
Plain RAG does not automatically understand that difference. It searches for text that appears similar and hopes the retrieved chunks are enough.
That fixed process is one reason for poor support chatbot accuracy. The model may write a polished reply, but the useful information never reached it.
The approach comes from a public n8n template called “Adaptive RAG strategy with query classification and retrieval”. My team studied the design and adapted it into a workflow using n8n, Google Gemini and Qdrant.
The workflow has five stages:
The adaptive RAG architecture adds another model call near the beginning. That call classifies the question and helps aim the later search.
Instead of asking one retrieval method to solve every customer question, the workflow provides different routes. The final retrieval and answer generation remain shared, which keeps the overall system easier to maintain.
Query classification gives the workflow a simple description of what the customer needs.
A Gemini model reads the question and returns one category name. The Switch node uses that exact word to choose a strategy route.
Example:
| Type | Example question | Search strategy |
|---|---|---|
| Factual | “What is the refund window for the annual plan?” | Make the search precise and narrow |
| Analytical | “How does plan A compare with plan B for a small clinic?” | Split the question into smaller searches |
| Opinion | “Is it worth moving from Excel to your tool?” | Search for different viewpoints in the business content |
| Contextual | “Will this work for my setup?” | Recover missing background and add it to the search |
The model must return one exact word.
A response such as “This appears to be an analytical question” sounds natural, but it can break the workflow. The Switch expects Analytical, not a complete sentence.
The classifier prompt follows the public template’s approach:
You classify one user question into exactly one category:
Factual, Analytical, Opinion or Contextual.
Factual = asks for a specific fact.
Analytical = asks for an explanation or comparison.
Opinion = asks for a view or recommendation.
Contextual = depends on the user's own situation.
Reply with the category name only.
Short or mixed questions can still confuse the classifier. Messages beginning with “Should I” may be Opinion or Contextual, while a message such as “pricing?” gives the model little information.
Hinglish also needs testing because Indian customers commonly mix English and Hindi in one sentence. I add three or four real examples for each difficult pattern so the prompt reflects how the product’s customers actually write.
Some questions fit more than one category. In that case, I send the Switch fallback to the Factual route because a precise search is usually the safest default.
Fallback questions are logged for review each week. This lets the business improve the classifier without rebuilding the rest of the workflow.
Factual questions need a narrow search that aims at one specific clause, number or rule.
Example:
The customer writes, “Refund?”
The Factual agent can rewrite that message as:
refund policy annual subscription cancellation window
The improved version is short but more specific. It has a better chance of retrieving the exact refund condition rather than a general billing page or unrelated mention of cancellation.
The Factual route should not widen the question or add opinions. Its job is to remove ambiguity while preserving the fact the customer wants.
The answer agent can then return a short response based on the retrieved clause. If the clause is not present, the bot should say it does not know rather than guess.
Analytical questions usually contain several hidden questions.
Example:
A customer asks, “Which plan suits a clinic with two doctors?”
Answering that may require information about user limits, appointment features, plan restrictions and price. Searching the complete sentence once may retrieve only one of those areas.
The Analytical agent separates the question into smaller searches. It retrieves content for the relevant parts, and the answer agent combines the findings into one comparison.
This route is useful for questions containing words such as compare, difference, better or why. It gives the retrieval step several clear targets instead of one broad sentence.
A comparison still has to stay within the business’s documents. The answer agent should not invent advantages that were never written in the source material.
An Opinion question should not return the first positive paragraph the vector search finds.
Example:
A customer asks, “Is it worth moving from Excel to your tool?”
The Opinion agent identifies angles such as cost, time saved and learning curve. It searches for information covering those different sides.
The final response can then explain the trade-offs using the business’s own content. It does not need to turn the answer into a sales pitch.
This route works only when the documents contain enough information to present more than one side. If every chunk says only that the product is excellent, the chatbot cannot produce a useful balanced answer without inventing details.
The better fix is improving the source documents, not writing a more forceful prompt.
Contextual questions depend on information the customer has already shared.
Example:
A customer asks, “Will this work for me?”
That sentence has almost no meaning by itself. The Contextual agent checks recent chat memory for useful background.
If the customer previously said they run a kirana store in Nellore, the agent can add that detail to the search. The improved query can then look for features, limitations or setup information relevant to that type of business.
The chat session ID connects messages belonging to the same customer. Without it, the workflow cannot know what “for me” refers to.
Context should come from the conversation, not from an unsupported guess. If the missing background is still unclear, the chatbot should ask the customer a follow-up question.
The complete public template contains around 39 nodes when its model and memory sub-nodes are included. The main path is shorter.
Information moves through these nodes in order:
This receives the customer’s question, chat session ID and collection ID. The workflow can be called from the chat interface, another workflow or a website integration.
This maps the three inputs into fixed field names. Later nodes always know where to find the question, session ID and collection ID.
The classifier returns Factual, Analytical, Opinion or Contextual.
The Switch reads the category and sends the question to the matching route. Its fallback connects to Factual.
Only one strategy agent runs for each question. It narrows, splits, widens or adds context to the original query.
The improved query becomes a vector and searches the selected collection. The node returns the closest document chunks.
The answer agent receives the retrieved chunks and the original question type. It generates a response using only the available content and a prompt designed for that category.
The final answer returns to the website, chat interface or calling workflow.
The final answer can only be as good as the chunks retrieved before it. A clever prompt cannot recover information the search never found.
At this point, a founder can judge if the extra routing is solving a real support problem or only making a small FAQ bot more complicated. Email upcomingtool@gmail.com with one direct customer question and one comparison question from your support history, and I will tell you where the routes would differ.
The Switch and strategy agents together form the RAG query router.
It sits before retrieval because Qdrant searches using the improved query it receives. A vague query is likely to return vague or incomplete chunks.
Once the wrong chunks have been retrieved, the answer agent has limited options. It cannot quote a plan feature, refund condition or eligibility rule that does not exist in its context.
The answer agent receives two inputs after retrieval:
The category helps control the response style. A factual question can receive a short answer, while an opinion question can present the different angles found in the content.
The RAG query router does not replace vector search. It prepares the customer’s question so the search has a better chance of retrieving useful evidence.
If the router is disconnected, every question loses the strategy chosen for its type. The workflow either stops or sends the unadapted question into retrieval.
The support documents must be loaded into Qdrant before the main chat workflow can answer from them.
I use a separate loader workflow. It reads the files, splits them into chunks, converts each chunk into a vector through the Embeddings Google Gemini node and inserts those vectors into a Qdrant collection.
One rule cannot be compromised: the loader and search workflow must use the same embedding model.
Vectors produced by different models do not match correctly. If the documents use one model and the search uses another, the Qdrant vector store may return irrelevant chunks even when the right answer exists in the source file.
A few loading rules help retrieval:
Large PDF-sized chunks can match many questions weakly without answering any of them properly. Smaller chunks give the search a clearer target.
The collection can run through Qdrant Cloud or a self-hosted Qdrant installation. Gemini requests still go to Google’s API in either setup.
Provider regions and terms can change, so check their current documentation before deciding where business documents should be processed. The UpcomingTools privacy policy explains how documents shared with my team during a build are handled.
I don’t know rule to the answer agent.Daily operation does not require writing code. Document updates and loader runs can be managed through the platform’s visual interface.
Prompt edits still need care. A small instruction can change the route or answer style for many customer questions.
Different failures can produce similar wrong answers, so the execution path needs to be checked in order.
| Failure | Result | Fix |
|---|---|---|
| Classifier returns a sentence | Switch sends questions to fallback | Require one category word and test with ten questions |
| Switch route is disconnected | That category stops before retrieval | Connect every output to its correct strategy agent |
| Wrong collection ID | Chatbot searches the wrong content or nothing | Confirm Combined Fields passes the expected ID |
| Qdrant collection is empty | Answer ignores business documents | Run the loader and confirm vectors exist |
| Embedding models do not match | Irrelevant chunks appear | Reload using the same Gemini embeddings model used for search |
| Strategy agent is disconnected | Improved query never reaches Qdrant | Reconnect the strategy to Retrieve Documents |
| Gemini quota is reached | Customer receives an error or no reply | Check the quota and add a friendly error response |
| Retrieval is disconnected from answer agent | Model receives no source chunks | Restore the document connection |
| Webhook responds immediately | Calling application gets an empty response | Respond through Respond to Webhook |
| Respond to Webhook is disconnected | The workflow generates an answer but the website receives nothing | Restore the final response connection |
Testing must cover direct questions, comparisons, recommendations, short messages, spelling mistakes and mixed-language inputs.
Neat questions written by the internal team are not enough. Customers will use incomplete sentences and product terms that do not match the documents exactly.
A few impressive answers are not a useful launch test.
My team collects fifty real questions from earlier support emails, WhatsApp chats or tickets. Each chatbot answer is marked correct, partial or wrong.
The scorecard shows whether the problem comes from query classification, missing documents, stale information, weak chunks or an answer prompt that ignores its evidence.
Money, refund and eligibility questions should move to a person when the chatbot cannot answer reliably from the retrieved content.
This gives support chatbot accuracy a visible measure. It is more useful than asking whether the bot sounds intelligent.
The answer agent also needs an I don’t know rule. When the retrieved chunks do not contain the answer, admitting that is safer than producing a confident guess.
The contact page can be used to share the types of questions your support team handles, with customer names and private account information removed.
| Point | Plain RAG | Adaptive RAG | Agentic RAG |
|---|---|---|---|
| Search process | One method for every question | One of several fixed strategies is selected | An agent chooses tools and how often to call them |
| Speed | Fastest | Adds a classification step | Slowest and can vary by question |
| Predictability | High | High | Lower |
| Useful for | Simple FAQ bots | Support and course-help bots with mixed questions | Research assistants using several data sources |
Adaptive retrieval occupies a practical middle position for many small Indian teams. It handles mixed questions without giving the model complete control over the workflow.
Agentic RAG becomes useful when questions repeatedly require several sources or tools, such as a CRM, a business database and a document collection.
It is harder to test, and the cost per question is less predictable. I move towards agentic RAG only when fixed routing clearly cannot handle the required support requests.
The node order can remain the same while the documents, metadata and safety rules change.
Coaching institutes: Separate collections can store fees, syllabus and schedules. The Factual route should be strict for fee questions.
Clinics: Document answers should stay limited to services, timings and bookings. Medical questions should go to staff.
SaaS support teams: Plan and version metadata should be attached to each chunk so answers match the customer’s subscription.
D2C brands: Collections can contain returns, delivery areas and product-care details. The Contextual route can use the customer’s city when that information is available.
The following terms help when discussing the build:
| Term | Meaning |
|---|---|
| Chunk | A small section of a document stored separately |
| Embedding | A list of numbers representing the meaning of a chunk |
| Vector store | A database that retrieves chunks with similar meaning |
| Top-k | The number of chunks returned by retrieval |
| RAG query router | The part that decides how a question should be searched |
| Session ID | The value connecting messages from one customer |
Old information needs active removal. A previous price list can be retrieved as easily as the current version if both remain in the collection.
Store a last-updated date in the metadata and remove stale chunks when policies, prices or conditions change.
| Stage | Work | What you receive |
|---|---|---|
| Audit | My team reviews the documents and earlier customer questions | A list of gaps in the source content |
| Load | My team splits and embeds documents in Qdrant | A searchable knowledge base |
| Tune | My team adjusts the classifier and strategy prompts | Prompts matching the product’s language |
| Test | My team runs the fifty-question scorecard | A correct, partial and wrong result report |
| Handover | My team connects the chatbot and explains updates | Access, notes and a recording |
From the first call to a working router usually takes my team 7 to 10 days, and it can take a little longer when your documents need cleaning first.
The exact shape depends on how many sources you have and which questions your users really ask, so we settle that together on calls and emails. I walk you through the design on a phone call or a Zoom call, and screen sharing lets you watch it come together.
The example in this article runs on one automation platform, but that is only a demonstration. If your team already works in another tool, or wants something built around your own systems, my team builds the same router there instead.
The Qdrant collection, Gemini credentials, documents and automation workflows remain on the client’s accounts.
The refund policy explains the terms for paid implementation work before the build starts.
An adaptive RAG architecture adds value only when the support questions need different search methods.
If the business has only twenty short FAQs, plain RAG or a clear FAQ page may be enough. Classification and four strategy agents would add work without solving a real problem.
If most questions concern a customer’s current orders or account, the chatbot needs access to a business database rather than only documents. That requires another type of workflow.
If the business has no written documentation, create it first. Retrieval cannot find an answer that has never been recorded.
Hindi and Telugu documents can be loaded, but the searches need testing in the same language. Confirm that the selected embedding model handles the language before making promises to customers.
Changing the embedding model later requires reloading the complete collection. Vectors made by the old model will not match the new ones correctly.
Choose simple RAG when the questions are direct and the answer normally lives in one clear chunk.
Choose adaptive RAG when customers ask for facts, comparisons, opinions and context from the same knowledge base.
Read about UpcomingTools and how my team approaches custom automation before deciding how much architecture your chatbot really needs.