The Adaptive RAG Architecture I Build for Indian SaaS and Edtech Founders’ Chatbots

The Adaptive RAG Architecture I Build for Indian SaaS and Edtech Founders’ Chatbots

The Adaptive RAG Architecture I Build for Indian SaaS and Edtech Founders’ Chatbots

Does your chatbot search for a direct fact, a comparison and a personal recommendation in exactly the same way? An adaptive RAG architecture classifies the question first, then chooses a search strategy that fits what the customer actually needs.

Basic RAG can answer direct questions well. It becomes less reliable when a customer asks the bot to compare plans, explain a choice or consider their particular situation.

The problem is not always the model. Sometimes the search method is wrong before the model even begins writing.

Why does a chatbot give a confident but wrong answer?

A basic RAG chatbot receives a question, searches for document chunks with similar meaning and sends the closest chunks to a language model.

That process works when the question has a direct answer in one place. It struggles when the answer depends on several documents, different viewpoints or context from an earlier message.

Example:

An online coaching platform loads its fee structure, course syllabus and refund rules into a basic RAG chatbot. A student asks, “What is the fee for the weekend batch?” and receives the correct amount from the fee document.

A parent then asks, “Is the weekend batch better than the regular batch for a working student?” The chatbot returns the same fee paragraph because it searches both questions through the same retrieval path.

The first question needs a precise lookup. The second needs a comparison across several pieces of content.

Plain RAG does not automatically understand that difference. It searches for text that appears similar and hopes the retrieved chunks are enough.

That fixed process is one reason for poor support chatbot accuracy. The model may write a polished reply, but the useful information never reached it.

How does an adaptive RAG architecture change the search?

The approach comes from a public n8n template called “Adaptive RAG strategy with query classification and retrieval”. My team studied the design and adapted it into a workflow using n8n, Google Gemini and Qdrant.

The workflow has five stages:

  1. Receive the question, chat session ID and Qdrant collection ID.
  2. Classify the question as Factual, Analytical, Opinion or Contextual.
  3. Adapt the question using a strategy for that category.
  4. Retrieve matching chunks from Qdrant with the improved query.
  5. Generate an answer using the original question type and retrieved content.

The adaptive RAG architecture adds another model call near the beginning. That call classifies the question and helps aim the later search.

Instead of asking one retrieval method to solve every customer question, the workflow provides different routes. The final retrieval and answer generation remain shared, which keeps the overall system easier to maintain.

What does query classification decide?

Query classification gives the workflow a simple description of what the customer needs.

A Gemini model reads the question and returns one category name. The Switch node uses that exact word to choose a strategy route.

Example:

TypeExample questionSearch strategy
Factual“What is the refund window for the annual plan?”Make the search precise and narrow
Analytical“How does plan A compare with plan B for a small clinic?”Split the question into smaller searches
Opinion“Is it worth moving from Excel to your tool?”Search for different viewpoints in the business content
Contextual“Will this work for my setup?”Recover missing background and add it to the search

The model must return one exact word.

A response such as “This appears to be an analytical question” sounds natural, but it can break the workflow. The Switch expects Analytical, not a complete sentence.

The classifier prompt follows the public template’s approach:

You classify one user question into exactly one category:
Factual, Analytical, Opinion or Contextual.
Factual = asks for a specific fact.
Analytical = asks for an explanation or comparison.
Opinion = asks for a view or recommendation.
Contextual = depends on the user's own situation.
Reply with the category name only.

Short or mixed questions can still confuse the classifier. Messages beginning with “Should I” may be Opinion or Contextual, while a message such as “pricing?” gives the model little information.

Hinglish also needs testing because Indian customers commonly mix English and Hindi in one sentence. I add three or four real examples for each difficult pattern so the prompt reflects how the product’s customers actually write.

Some questions fit more than one category. In that case, I send the Switch fallback to the Factual route because a precise search is usually the safest default.

Fallback questions are logged for review each week. This lets the business improve the classifier without rebuilding the rest of the workflow.

How are factual questions searched?

Factual questions need a narrow search that aims at one specific clause, number or rule.

Example:

The customer writes, “Refund?”

The Factual agent can rewrite that message as:

refund policy annual subscription cancellation window

The improved version is short but more specific. It has a better chance of retrieving the exact refund condition rather than a general billing page or unrelated mention of cancellation.

The Factual route should not widen the question or add opinions. Its job is to remove ambiguity while preserving the fact the customer wants.

The answer agent can then return a short response based on the retrieved clause. If the clause is not present, the bot should say it does not know rather than guess.

How are analytical questions broken down?

Analytical questions usually contain several hidden questions.

Example:

A customer asks, “Which plan suits a clinic with two doctors?”

Answering that may require information about user limits, appointment features, plan restrictions and price. Searching the complete sentence once may retrieve only one of those areas.

The Analytical agent separates the question into smaller searches. It retrieves content for the relevant parts, and the answer agent combines the findings into one comparison.

This route is useful for questions containing words such as compare, difference, better or why. It gives the retrieval step several clear targets instead of one broad sentence.

A comparison still has to stay within the business’s documents. The answer agent should not invent advantages that were never written in the source material.

How are opinion questions kept balanced?

An Opinion question should not return the first positive paragraph the vector search finds.

Example:

A customer asks, “Is it worth moving from Excel to your tool?”

The Opinion agent identifies angles such as cost, time saved and learning curve. It searches for information covering those different sides.

The final response can then explain the trade-offs using the business’s own content. It does not need to turn the answer into a sales pitch.

This route works only when the documents contain enough information to present more than one side. If every chunk says only that the product is excellent, the chatbot cannot produce a useful balanced answer without inventing details.

The better fix is improving the source documents, not writing a more forceful prompt.

How do contextual questions recover missing information?

Contextual questions depend on information the customer has already shared.

Example:

A customer asks, “Will this work for me?”

That sentence has almost no meaning by itself. The Contextual agent checks recent chat memory for useful background.

If the customer previously said they run a kirana store in Nellore, the agent can add that detail to the search. The improved query can then look for features, limitations or setup information relevant to that type of business.

The chat session ID connects messages belonging to the same customer. Without it, the workflow cannot know what “for me” refers to.

Context should come from the conversation, not from an unsupported guess. If the missing background is still unclear, the chatbot should ask the customer a follow-up question.

How does every node move the question forward?

The complete public template contains around 39 nodes when its model and memory sub-nodes are included. The main path is shorter.

Information moves through these nodes in order:

  1. Chat Trigger or Execute Workflow Trigger

This receives the customer’s question, chat session ID and collection ID. The workflow can be called from the chat interface, another workflow or a website integration.

  1. Combined Fields (Edit Fields)

This maps the three inputs into fixed field names. Later nodes always know where to find the question, session ID and collection ID.

  1. Query Classification (AI Agent + Gemini)

The classifier returns Factual, Analytical, Opinion or Contextual.

  1. Switch

The Switch reads the category and sends the question to the matching route. Its fallback connects to Factual.

  1. Four strategy agents (Gemini)

Only one strategy agent runs for each question. It narrows, splits, widens or adds context to the original query.

  1. Retrieve Documents (Qdrant Vector Store + Gemini embeddings)

The improved query becomes a vector and searches the selected collection. The node returns the closest document chunks.

  1. Answer agent (Gemini + chat memory)

The answer agent receives the retrieved chunks and the original question type. It generates a response using only the available content and a prompt designed for that category.

  1. Respond to Webhook

The final answer returns to the website, chat interface or calling workflow.

The final answer can only be as good as the chunks retrieved before it. A clever prompt cannot recover information the search never found.

At this point, a founder can judge if the extra routing is solving a real support problem or only making a small FAQ bot more complicated. Email upcomingtool@gmail.com with one direct customer question and one comparison question from your support history, and I will tell you where the routes would differ.

Why must the RAG query router come before retrieval?

The Switch and strategy agents together form the RAG query router.

It sits before retrieval because Qdrant searches using the improved query it receives. A vague query is likely to return vague or incomplete chunks.

Once the wrong chunks have been retrieved, the answer agent has limited options. It cannot quote a plan feature, refund condition or eligibility rule that does not exist in its context.

The answer agent receives two inputs after retrieval:

  • The document chunks.
  • The original question category.

The category helps control the response style. A factual question can receive a short answer, while an opinion question can present the different angles found in the content.

The RAG query router does not replace vector search. It prepares the customer’s question so the search has a better chance of retrieving useful evidence.

If the router is disconnected, every question loses the strategy chosen for its type. The workflow either stops or sends the unadapted question into retrieval.

How do the Qdrant vector store and Gemini embeddings work?

The support documents must be loaded into Qdrant before the main chat workflow can answer from them.

I use a separate loader workflow. It reads the files, splits them into chunks, converts each chunk into a vector through the Embeddings Google Gemini node and inserts those vectors into a Qdrant collection.

One rule cannot be compromised: the loader and search workflow must use the same embedding model.

Vectors produced by different models do not match correctly. If the documents use one model and the search uses another, the Qdrant vector store may return irrelevant chunks even when the right answer exists in the source file.

A few loading rules help retrieval:

  • Keep each chunk focused on one clear idea, such as one policy clause or FAQ answer.
  • Add metadata including product, plan, language and last-updated date.
  • Pass the collection ID into the workflow so the same architecture can support several products or businesses.
  • Remove old chunks when prices, policies or eligibility conditions change.
  • Use the same Gemini embeddings model for insertion and retrieval.

Large PDF-sized chunks can match many questions weakly without answering any of them properly. Smaller chunks give the search a clearer target.

The collection can run through Qdrant Cloud or a self-hosted Qdrant installation. Gemini requests still go to Google’s API in either setup.

Provider regions and terms can change, so check their current documentation before deciding where business documents should be processed. The UpcomingTools privacy policy explains how documents shared with my team during a build are handled.

Build order inside the workflow

  1. Create a Qdrant collection through Qdrant Cloud or a self-hosted installation.
  2. Add the Qdrant credential to the workflow platform.
  3. Add a Google Gemini API credential.
  4. Build or import the loader workflow.
  5. Load the help documents into the collection.
  6. Import the adaptive workflow.
  7. Confirm the input names inside Combined Fields.
  8. Replace the classifier and strategy examples with real product questions.
  9. Point Retrieve Documents to the correct collection.
  10. Set how many chunks retrieval should return.
  11. Add the I don’t know rule to the answer agent.
  12. Connect the workflow to the n8n chat interface or website widget.
  13. Test fifty real questions and mark each result correct, partial or wrong.

Daily operation does not require writing code. Document updates and loader runs can be managed through the platform’s visual interface.

Prompt edits still need care. A small instruction can change the route or answer style for many customer questions.

What breaks when a node fails?

Different failures can produce similar wrong answers, so the execution path needs to be checked in order.

FailureResultFix
Classifier returns a sentenceSwitch sends questions to fallbackRequire one category word and test with ten questions
Switch route is disconnectedThat category stops before retrievalConnect every output to its correct strategy agent
Wrong collection IDChatbot searches the wrong content or nothingConfirm Combined Fields passes the expected ID
Qdrant collection is emptyAnswer ignores business documentsRun the loader and confirm vectors exist
Embedding models do not matchIrrelevant chunks appearReload using the same Gemini embeddings model used for search
Strategy agent is disconnectedImproved query never reaches QdrantReconnect the strategy to Retrieve Documents
Gemini quota is reachedCustomer receives an error or no replyCheck the quota and add a friendly error response
Retrieval is disconnected from answer agentModel receives no source chunksRestore the document connection
Webhook responds immediatelyCalling application gets an empty responseRespond through Respond to Webhook
Respond to Webhook is disconnectedThe workflow generates an answer but the website receives nothingRestore the final response connection

Testing must cover direct questions, comparisons, recommendations, short messages, spelling mistakes and mixed-language inputs.

Neat questions written by the internal team are not enough. Customers will use incomplete sentences and product terms that do not match the documents exactly.

How do you measure support chatbot accuracy?

A few impressive answers are not a useful launch test.

My team collects fifty real questions from earlier support emails, WhatsApp chats or tickets. Each chatbot answer is marked correct, partial or wrong.

The scorecard shows whether the problem comes from query classification, missing documents, stale information, weak chunks or an answer prompt that ignores its evidence.

Money, refund and eligibility questions should move to a person when the chatbot cannot answer reliably from the retrieved content.

This gives support chatbot accuracy a visible measure. It is more useful than asking whether the bot sounds intelligent.

The answer agent also needs an I don’t know rule. When the retrieved chunks do not contain the answer, admitting that is safer than producing a confident guess.

The contact page can be used to share the types of questions your support team handles, with customer names and private account information removed.

How are adaptive RAG and agentic RAG different?

PointPlain RAGAdaptive RAGAgentic RAG
Search processOne method for every questionOne of several fixed strategies is selectedAn agent chooses tools and how often to call them
SpeedFastestAdds a classification stepSlowest and can vary by question
PredictabilityHighHighLower
Useful forSimple FAQ botsSupport and course-help bots with mixed questionsResearch assistants using several data sources

Adaptive retrieval occupies a practical middle position for many small Indian teams. It handles mixed questions without giving the model complete control over the workflow.

Agentic RAG becomes useful when questions repeatedly require several sources or tools, such as a CRM, a business database and a document collection.

It is harder to test, and the cost per question is less predictable. I move towards agentic RAG only when fixed routing clearly cannot handle the required support requests.

How can Indian businesses customise the routes?

The node order can remain the same while the documents, metadata and safety rules change.

Coaching institutes: Separate collections can store fees, syllabus and schedules. The Factual route should be strict for fee questions.

Clinics: Document answers should stay limited to services, timings and bookings. Medical questions should go to staff.

SaaS support teams: Plan and version metadata should be attached to each chunk so answers match the customer’s subscription.

D2C brands: Collections can contain returns, delivery areas and product-care details. The Contextual route can use the customer’s city when that information is available.

The following terms help when discussing the build:

TermMeaning
ChunkA small section of a document stored separately
EmbeddingA list of numbers representing the meaning of a chunk
Vector storeA database that retrieves chunks with similar meaning
Top-kThe number of chunks returned by retrieval
RAG query routerThe part that decides how a question should be searched
Session IDThe value connecting messages from one customer

Old information needs active removal. A previous price list can be retrieved as easily as the current version if both remain in the collection.

Store a last-updated date in the metadata and remove stale chunks when policies, prices or conditions change.

How does my team deliver the workflow?

StageWorkWhat you receive
AuditMy team reviews the documents and earlier customer questionsA list of gaps in the source content
LoadMy team splits and embeds documents in QdrantA searchable knowledge base
TuneMy team adjusts the classifier and strategy promptsPrompts matching the product’s language
TestMy team runs the fifty-question scorecardA correct, partial and wrong result report
HandoverMy team connects the chatbot and explains updatesAccess, notes and a recording

From the first call to a working router usually takes my team 7 to 10 days, and it can take a little longer when your documents need cleaning first.

The exact shape depends on how many sources you have and which questions your users really ask, so we settle that together on calls and emails. I walk you through the design on a phone call or a Zoom call, and screen sharing lets you watch it come together.

The example in this article runs on one automation platform, but that is only a demonstration. If your team already works in another tool, or wants something built around your own systems, my team builds the same router there instead.

The Qdrant collection, Gemini credentials, documents and automation workflows remain on the client’s accounts.

The refund policy explains the terms for paid implementation work before the build starts.

When is an adaptive RAG architecture unnecessary?

An adaptive RAG architecture adds value only when the support questions need different search methods.

If the business has only twenty short FAQs, plain RAG or a clear FAQ page may be enough. Classification and four strategy agents would add work without solving a real problem.

If most questions concern a customer’s current orders or account, the chatbot needs access to a business database rather than only documents. That requires another type of workflow.

If the business has no written documentation, create it first. Retrieval cannot find an answer that has never been recorded.

Hindi and Telugu documents can be loaded, but the searches need testing in the same language. Confirm that the selected embedding model handles the language before making promises to customers.

Changing the embedding model later requires reloading the complete collection. Vectors made by the old model will not match the new ones correctly.

Which search setup should you choose?

Choose simple RAG when the questions are direct and the answer normally lives in one clear chunk.

Choose adaptive RAG when customers ask for facts, comparisons, opinions and context from the same knowledge base.

Read about UpcomingTools and how my team approaches custom automation before deciding how much architecture your chatbot really needs.

Tags
Share Article:

Radha Krishna

Leave a Comment

UpcomingTools logo - Nation First, Build The Best, Pass The Test