Tensic guides

Guide 12

Build a RAG knowledge base and connect it to an agent

Create a RAG project, ingest documents or websites, tune the score cutoff and Top-K so it only answers from your content, and let agents search it with the search_knowledge tool.

On this page
  1. Create a RAG project
  2. Ingest documents into a knowledge base
  3. Test retrieval in the Knowledge workbench
  4. Set a score cutoff to fence a RAG project
  5. Choose Top-K and LLM rerank
  6. Ingest a website and keep it in sync
  7. Connect an agent to a knowledge base

A RAG (retrieval-augmented generation) project answers questions from your own documents. You ingest files or websites into its knowledge base, and the project retrieves the most relevant chunks for each question and answers from them. A RAG project can be called directly through its API, or used as a knowledge source by an agent.

Create a RAG project

  1. Go to Projects and create a new project.
  2. Enter a Project name and select the Team.
  3. Under Type, select RAG ("Answer questions from documents you upload"). The type is fixed once the project is created; a project cannot be converted to another type later.
  4. Select the LLM, the model that writes the answers.
  5. Select the Embeddings model, the vector encoder used for retrieval.
  6. Select the Vector store, where the embedded chunks are stored. The stores offered depend on your Tensic instance; PGVector is one option.
  7. Check the Live spec summary on the right and click Create project.
New Project form with type RAG selected and the LLM, Embeddings and Vector store fields

RAG projects have a Knowledge tab, where you manage the documents. Agent projects do not have this tab. They use a RAG project as their knowledge instead (see "Connect an agent to a knowledge base").

Ingest documents into a knowledge base

  1. Open the RAG project and go to the Knowledge tab.
  2. Under Data, stay on File and drop your files, or click to browse. Use the URL tab to ingest a single web page once.
  3. Click Ingest.
  4. The file appears in the Ingest Queue. All ingestion goes through a queue, so you can load thousands of documents (for example overnight) as well as a single file, which is processed within seconds. When the status shows done, the queue entry shows the number of chunks created.
  5. Under Documents, expand a document to see its chunks and the extracted text.

The header of the Knowledge tab shows the totals: Documents, Chunks, the Embeddings model and the Vector store. Delete a document with the bin icon next to it.

Test retrieval in the Knowledge workbench

The Workbench on the Knowledge tab shows exactly what retrieval returns for a query, before you expose the project to users.

  1. In Query, type a question.
  2. Set Top-K and Cutoff for this test run and click Run.
  3. Retrieved โ€” raw similarity lists the chunks found, with a similarity score between 0 and 1 for each.
  4. Answer โ€” full pipeline shows the answer the project gives a real caller with those settings.
Knowledge workbench: a query unrelated to the document retrieves a chunk with a low score of 0.28

A low score means the chunk is not really related to the question. With a cutoff of 0, the project still passes that chunk to the model, and the model may answer from its own general knowledge instead of from your documents. The answer can then be outdated or invented. The next section shows how to stop that.

Set a score cutoff to fence a RAG project

The Score cutoff is the most important RAG setting. It is the minimum similarity a chunk needs to be used. When no chunk reaches it, the project does not answer from general model knowledge; it returns the Default fallback answer instead. Setting it is called fencing the knowledge base.

Find the right cutoff by testing:

  1. In the workbench, ask several questions you know are not covered by your documents, with Cutoff set to 0. Note the scores. In the example, an unrelated question about the company scored 0.28 against an invoice.
  2. Ask questions that are covered and note their scores. In the example, a name printed on the invoice scored 0.36.
  3. Choose a cutoff above the unrelated scores and below the relevant ones, for example 0.30.
  4. Under Retrieval, enter it in Score cutoff (0โ€“1; higher is stricter) and click Save.
  5. Open the Guards tab and set the Default fallback answer, the message returned when no chunk meets the cutoff (for example "I don't know the answer to that. Please contact support.").
  6. Re-test in the Playground: small talk such as "hi" and off-topic questions get the fallback answer, while questions about the document are answered.
Retrieval settings: Top-K slider, Score cutoff 0.30 and the LLM rerank switch

Fencing cannot be automated reliably, because only you know what is and is not in your documents. Repeat the test whenever you add a lot of new content. For anything customer-facing, such as a website widget, always set a cutoff.

Choose Top-K and LLM rerank

  • Top-K (chunks retrieved per query) sets how many of the highest-scoring chunks are used for each question. A higher Top-K gives the model more context and usually better answers, but each extra chunk adds processing, so answers take longer and cost more. Start with a small value such as 4 and change it only when the evals or tests show a need.
  • LLM rerank re-scores the retrieved chunks with an LLM before they are used. This is more accurate, at the cost of one extra model call per query. Reranking applies to the full-pipeline answer, not to the raw similarity scores in the workbench.

Measure the effect of any retrieval change with an eval. RAG projects offer the extra Faithfulness metric. See Run evals and manage prompt versions.

Ingest a website and keep it in sync

To index a website and keep the index current, add it as a sync source:

  1. On the Knowledge tab, go to Sources and click + Add source.
  2. Enter a Name for the source, for example news.
  3. Set Type to Web URL and enter the URL.
  4. Choose the Sync interval, for example every 15 minutes.
  5. Choose the Splitter (for example Sentence) and Chunk size (for example 512).
  6. Click Save. Saving commits only this source; the retrieval settings are not changed.
  7. Turn on Auto-sync so each source runs on its interval, or click Sync now to run it immediately.
  8. If the site needs a login, click Credentials on the source and enter them. Credentials are matched to the source by name, so create a new source rather than renaming one.

New pages appear under Documents after each sync. To be notified when a sync finishes, subscribe to the sync_completed event under Integrations > Event webhooks.

Add source dialog with type Web URL, a 15-minute sync interval, sentence splitter and chunk size 512

For a one-time import without syncing, use the URL tab under Data instead.

Connect an agent to a knowledge base

An agent can use one RAG project as its knowledge base. The agent sends its search to the RAG project, which applies its own Top-K, score cutoff and fallback behaviour. This setup is called agentic RAG: the agent decides when to search, and the RAG project handles retrieval.

  1. Open the agent project and go to the Tools & MCP tab.
  2. Under Behaviour > Built-in tools, add search_knowledge and save.
  3. Go to the Orchestration tab. In the wiring diagram, click the Knowledge node (add knowledge).
  4. Under Knowledge edge, select your RAG project in Knowledge source and click Save.
Tools & MCP tab with search_knowledge in the built-in tools
Orchestration tab with the Knowledge node and the Knowledge source field
  1. Tell the agent in its prompt what the knowledge base contains and when to use it, for example: "You have access to a knowledge base with invoices. Use the search_knowledge tool to answer questions about them." An agent with many tools may otherwise not know when to search.
  2. Test in the agent's Playground with a question that only the documents can answer.

You can combine the knowledge base with other built-in tools, such as web search, on the same agent, so it can answer from both your documents and the web.

Common questions

Why does my RAG project answer questions that are not in my documents?

The score cutoff is too low (0 by default), so weak matches are used and the model fills in with general knowledge. Raise the Score cutoff and set a Default fallback answer.

How do I find the right score cutoff?

Ask questions you know are not in your documents and note their scores. Then set the cutoff just above the highest of those scores, below the scores of real matches.

Does a higher Top-K give better answers?

Usually, yes, but answers take longer and cost more. Start with a small value such as 4 and adjust based on tests.

Can I add a whole website?

Yes. Add a Web URL source with a Sync interval and turn on Auto-sync. New pages are indexed on every run.

How long does ingestion take?

Files go through a queue. A single file is processed within seconds, and large batches are processed in the background.

Do agents have a Knowledge tab?

No. Create a RAG project for the documents, add the search_knowledge tool to the agent, and select the RAG project as Knowledge source on the agent's Orchestration tab.

Can I change a project's type to RAG later?

No. The project type is fixed at creation. Create a new RAG project instead.