Skip to main content

How retrieval works

An agent working on a task often needs knowledge it does not have. Retrieval is how it gets that context. The agent sends a query, and Kapa searches the index that ingestion built from your knowledge sources and returns exactly the parts that are relevant to it: a section of a page, a post, a table, a block of code, each with the URL it came from. The agent then uses that context to carry on with its task, whether that is answering a user in your product, drafting a reply to a support ticket, or running a workflow. The Examples show retrieval wired into agents of each kind.

What makes a retriever good​

A good retriever has two qualities. It finds everything in your knowledge sources that the agent needs for the query: this is recall. And it does not also return information that is irrelevant to the query: this is precision.

Recall matters most. If the passage that answers the query is not returned, the agent never sees it, and the answer comes out incomplete or wrong. Nothing the agent does afterwards can fix that.

Precision matters too. Everything that is returned goes into the agent's context. Each unneeded passage costs tokens on every call, and as the context fills with material that has nothing to do with the query, the model reasons worse, a problem known as context rot.

The ideal retriever returns everything the agent needs and nothing else.

Kapa's retrieval has two modes:

  • deep is the most capable one, built to score high on both measures, and most of this page describes how it works.
  • default is the faster mode, described at the end of this page. Both are selected with the mode parameter on the HTTP API and the MCP server. To pick the right mode for your use case, see Choose a retrieval mode.

The deep mode​

The deep mode is not a single search over your knowledge sources. It is an agent of its own. It plans and uses its tools until it is confident that it has gathered everything needed to address the query, as far as your sources contain it. Along the way it builds its own understanding of your knowledge base. Addressing a query rarely takes a single piece of information. It usually takes several, often of different kinds and from different places in your sources, and the deep mode has to explore several concepts before it knows what to search for. It then returns what it judged necessary and leaves out the rest. This is how it scores well on both measures: gathering until it is confident gives it recall, and returning only what it has read and judged relevant gives it precision.

Tools​

The deep mode's tools give it the different capabilities it needs to find different kinds of information efficiently while it searches. Each is a fast, direct lookup against your index.

Knowledge base topology​

At the start, the deep mode accesses a topology of everything you have indexed: a map of what your sources cover, how the main areas relate, and the terminology they use. Kapa builds this topology from the sources themselves and keeps it up-to-date as they change, so the deep mode never starts a query without knowing the shape of the knowledge it is searching.

Searches your sources for passages that contain specific words. Rare words count for much more than common ones, so an error code, a configuration key, or a command name leads straight to the passages that mention it.

Searches your sources by meaning, using embeddings from a dense, multilingual text embedding model.

Rerank​

Orders a set of search results by how well each passage answers the query, using a cross-encoder: a model that reads the query and a passage together and scores the pair. The deep mode can use it to sort intermediate result sets for efficiency and keep only the most relevant passages.

Read document​

Reads more of a document the deep mode has already found, for example the rest of a long table, the remaining steps of a procedure, or the rest of a list.

Prune​

Drops passages that turned out not to be relevant to the query. The deep mode has read everything it gathered, so it can judge which passages the agent needs and return only those.

Budgets​

The search is bounded three ways: a maximum number of rounds, a time limit, and a ceiling on how many passages the context may hold. Reaching any of them ends the search with what has been gathered so far. Most queries take about two rounds; roughly one run in a hundred reaches a budget.

What it returns​

The deep mode returns a variable number of passages, depending on how much information in your sources is relevant to the query. A query answered by one passage returns one; a query that needs a whole table returns the table. top_k and max_chars still apply as an upper bound, but the deep mode rarely reaches it.

The default mode​

The default mode is designed to be the faster of the two modes. It does only limited planning and reasoning, and it has a condensed toolset compared to the deep mode.

Tools​

The default mode works with the same tools as the deep mode, excluding pruning and reading full documents, because those introduce more latency and require more reasoning.

Budgets​

The default mode's budget is much stricter than the deep mode's and does not allow for the same depth, which is what keeps it fast.

What it returns​

A fixed number of passages, the best ranked ones, up to top_k and within max_chars. The default mode returns that many whether or not all of them are relevant to the query.