LangChain
Kapa has a native LangChain integration, the langchain-kapa-ai Python package, that gives your LangChain chains and agents access to Kapa's retrieval over the knowledge sources in your Kapa project. It offers the same functionality as the hosted MCP server and the HTTP API, and it is the most native way to use Kapa's retrieval inside LangChain. Its source code, examples, and changelog are in the kapa-ai/langchain-kapa-ai repository on GitHub.
For a full walkthrough of installing the package and using it in an agent, see Add knowledge base search to a LangChain agent.
Requirements
- Python 3.10 or later.
langchain-core1.4 or later, before 2.0.- A Kapa project with indexed knowledge sources. To set one up, follow Index your first source.
Installation
Install the package from PyPI:
pip install langchain-kapa-ai
The import name is langchain_kapa_ai.
Components
The package provides three components:
KapaRetriever: returns the most relevant chunks for a query as LangChain documents.KapaGetDocumentsTool: an agent tool that fetches whole documents by source link or document ID.KapaToolkit: returns search and document lookup as agent tools, configured from one set of settings.
The examples below read the API key and project ID from the environment; Connection settings lists all settings.
KapaRetriever
KapaRetriever searches all knowledge sources connected to your Kapa project for a query and returns the most relevant chunks, each with the URL of its source. It is Kapa's retrieval, the same one the hosted MCP server exposes as its search tool.
The retriever extends LangChain's BaseRetriever, supports invoke, ainvoke, batch, and abatch, and calls the Retrieval endpoint, whose reference page defines the server-side defaults, limits, and fields.
You can configure the following parameters on the retriever:
| Setting | Type | Default | Description |
|---|---|---|---|
mode | "default" or "deep" | "default" | default is faster; deep searches further and returns only relevant chunks. See Choose a retrieval mode. |
top_k | int (1-15) | 15 | The maximum number of chunks to return. Fewer may be returned if max_chars or the deep mode reduce the result set. |
max_chars | int (1-60000) | 35,000 | Maximum number of characters across all returned chunks. Chunks are included in order, but only up to the point where the total character count stays within this limit. Chunks are never truncated. This is an upper bound, not a target: especially in the deep mode, the returned total may be well below this limit. |
source_group_ids | list[str] | All sources | Only return results from sources in these groups. |
integration_id | str | None | Integration that analytics attributes queries to. |
redact_query | bool | False | If True, the query text is redacted from analytics. Use for sensitive queries. |
end_user | KapaEndUser | None | Associates queries with an end user in your analytics: email (the user's email address), unique_client_id (your own identifier for the user), and metadata (company_name, first_name, and last_name; other keys are ignored). |
mode, top_k, and max_chars can also be overridden for a single call, for example retriever.invoke(query, top_k=3). Other keyword arguments on invoke raise TypeError.
from langchain_kapa_ai import KapaRetriever
retriever = KapaRetriever()
for document in retriever.invoke("How do I rotate an API key?", top_k=3):
print(document.metadata["source"])
print(document.page_content)
Results
The retriever returns a list of LangChain Document objects, best match first. Each document's page_content is the chunk text in Markdown, and metadata["source"] is the URL of its source, exactly as Kapa returns it. The URL can point inside the document, for example to a heading, a line range, or a PDF page. A call can return fewer chunks than top_k, or none.
To show the source URLs to a model, include {source} in the document prompt of LangChain's create_retriever_tool, which otherwise passes only the chunk text.
KapaGetDocumentsTool
KapaGetDocumentsTool is an agent tool that fetches whole documents from your knowledge sources by their source URL or document ID. An agent uses it when a chunk from search is not enough and it needs the complete page.
The tool extends LangChain's BaseTool, is named get_knowledge_documents, and calls the Documents endpoint, whose reference page defines the server-side defaults, limits, and fields. Credentials, the project, and source groups stay in application settings, so the agent cannot change them.
You can configure the following parameters on the tool:
| Setting | Type | Default | Description |
|---|---|---|---|
source_group_ids | list[str] | All sources | Only return documents from sources in these groups. |
max_chars_per_document | int (1-200000) | 50,000 | Maximum number of characters returned per document. Longer documents are truncated and flagged with truncated. |
page_size | int, minimum 1 | 5 | The number of requested URLs and document IDs per page. The tool splits a page into as many endpoint requests as needed. |
The agent passes these arguments on each call:
| Argument | Type | Description |
|---|---|---|
urls | list[str] | Source URLs as search results cite them. |
document_ids | list[str] | The IDs of the documents to fetch. |
page | int | The page of requested documents to return, starting at 1. Defaults to 1. |
At least one URL or document ID is required. Duplicate URLs and IDs are ignored. A URL or ID that matches no document has no entry in documents.
from langchain_kapa_ai import KapaGetDocumentsTool
tool = KapaGetDocumentsTool()
message = tool.invoke(
{
"name": "get_knowledge_documents",
"args": {"urls": ["https://docs.example.com/guide"]},
"id": "call-1",
"type": "tool_call",
}
)
message.content # JSON text for the model
page = message.artifact # KapaDocumentsPage for your code
for document in page.documents:
print(document.title, document.source_url)
Results
The tool returns a LangChain ToolMessage. Its content is a JSON string for the model, and its artifact is a KapaDocumentsPage object for your code, which the model does not see. Both carry the same fields:
| Field | Description |
|---|---|
page, page_size | The returned page and the page size. |
total_requested | The number of distinct requested URLs and document IDs. |
has_more, next_page | Whether another page of requested URLs and IDs exists, and its number. |
documents | Each document found on this page, once, as a KapaDocument, in request order. |
Each KapaDocument has these fields:
| Field | Description |
|---|---|
document_id | The ID of the document. |
source_url | The URL of the document, or None if it has no URL. |
title | The title of the document. |
content | The document in Markdown, truncated to max_chars_per_document, or None when the content is unavailable, such as for PDFs. |
content_available | Whether the document text is available. |
total_chars | The length of the full document before truncation, or 0 when the content is unavailable. |
truncated | Whether content was truncated to max_chars_per_document. |
KapaToolkit
KapaToolkit gives an agent both search and document lookup in one step. It creates a KapaRetriever and a KapaGetDocumentsTool from one set of parameters and returns them as two agent tools.
The toolkit extends LangChain's BaseToolkit. A retriever on its own is not an agent tool, so the toolkit wraps the KapaRetriever in a search tool with LangChain's create_retriever_tool. The KapaGetDocumentsTool is already a tool and is returned unchanged.
You can configure the following parameters on the toolkit: the connection settings, every KapaRetriever parameter, and every KapaGetDocumentsTool parameter. The toolkit passes each parameter to the tool that uses it, and source_group_ids to both.
This example also needs langchain and your model provider's LangChain package. Replace <provider>:<model> with a tool-capable model and configure the provider's credentials.
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain_kapa_ai import KapaToolkit
agent = create_agent(
init_chat_model("<provider>:<model>"),
tools=KapaToolkit().get_tools(),
system_prompt="Search the knowledge base before answering and cite your sources.",
)
Results
get_tools() returns a list of two tools:
| Tool | Name | Built from | Returns |
|---|---|---|---|
| Search tool | search_knowledge_sources | A KapaRetriever, wrapped with create_retriever_tool | A ToolMessage whose content lists the chunks as text, each as Source: {source} followed by the chunk text, and whose artifact is the list of LangChain Document objects described in KapaRetriever results. |
| Document tool | get_knowledge_documents | A KapaGetDocumentsTool, unchanged | The ToolMessage described in KapaGetDocumentsTool results. |
The model reads each tool's description to decide when to call it. The search tool's description explains what a chunk is and that chunks come back in order of relevance; the document tool's description says when to look up a whole document.
Connection settings
All three components accept these settings. By default, they read the API key and project ID from the environment:
export KAPA_API_KEY="your-api-key"
export KAPA_PROJECT_ID="your-project-id"
Pass a setting to override its default, for example to read the key from another variable or to search a different project:
import os
from langchain_kapa_ai import KapaRetriever
retriever = KapaRetriever(
api_key=os.environ["SUPPORT_KAPA_API_KEY"],
project_id="your-project-id",
timeout=30.0,
)
| Setting | Type | Default | Description |
|---|---|---|---|
api_key | str | KAPA_API_KEY environment variable | Project API key, sent in the X-API-KEY header. It is never shown in representations or error messages. |
project_id | str | KAPA_PROJECT_ID environment variable | The project to search. |
base_url | str | https://api.kapa.ai | API base URL. |
timeout | float | 60.0 | Timeout in seconds for each request. |
http_client | httpx.Client | None | Client for synchronous requests. Without one, each call opens and closes its own client. |
http_async_client | httpx.AsyncClient | None | Client for asynchronous requests, with the same default. |
No component retries a failed request. In a chain, wrap a component with LangChain's with_retry() to retry. In an agent, retry tool calls with LangChain's ToolRetryMiddleware. To turn a KapaError into a message the model can read, use wrap_tool_call.
Errors
Programming errors are raised before any request is sent, as standard Python exceptions: a missing API key or project ID raises ValueError, invalid settings or tool input raise pydantic's ValidationError, and an unsupported per-call argument raises TypeError. A value outside Kapa's limits raises KapaValidationError instead.
A failed request raises an exception that derives from KapaError; it never becomes an empty result. Exceptions for an error status derive from KapaAPIError, which carries status_code and the server's detail.
| Exception | Raised when |
|---|---|
KapaAuthenticationError | The server rejects the API key or its access to the project (HTTP 401 or 403). |
KapaNotFoundError | The project, or the integration in integration_id, does not exist for the key (HTTP 404). |
KapaValidationError | Kapa rejects the request parameters (other HTTP 4xx). |
KapaRateLimitError | A rate limit or quota is exceeded (HTTP 429). retry_after holds the server's wait time in seconds, when given. |
KapaServiceError | Kapa fails to process the request (HTTP 5xx). |
KapaConnectionError | The request does not reach Kapa or times out. |
KapaResponseError | The response does not match the expected format. |
Rate limits
The package shares the rate limits of the Retrieval and Documents endpoints, which apply per team across all projects and integrations.
Related
- Add knowledge base search to a LangChain agent: build an agent that combines native tools with search and document lookup.
- Choose a retrieval mode: when to set
modetodeep. - Prompt for grounded answers: tell the agent when to search and how to use the results.