Knowledge bases
Give agents external data and attach it where it is needed.
A knowledge base is an external data source the model can query while writing: scientific literature, encyclopedic entries, or your own catalog of offers, events, or recordings. They live under Library → Knowledge bases.
Knowledge bases never create articles or clusters. Nothing from them enters the Stream and nothing passes through relevance filtering — they only enrich the text being written.
Bases every account starts with
Two default bases are created for every account at registration:
- PubMed — abstracts of peer-reviewed publications from NCBI. It searches the last 365 days and returns up to 20 results by default.
- Wikipedia (pl) — Polish Wikipedia entries: definitions, background, and conceptual context.
Both run in live mode: the external API is queried only when the model asks for something. Nothing is stored on our side and nothing is billed until they are used, which is why they are available on every plan.
Live mode and indexed mode
| Live | Indexed | |
|---|---|---|
| When the source is queried | at the moment the model calls it | ahead of time, on a schedule |
| Where the data lives | nowhere — the response is used once | in SemanticHub, as entries with embeddings |
| Retrieval | whatever the external API offers | semantic, by content similarity |
| Cost | only on use | per new and changed entry |
Default bases are live. A base you add yourself is indexed: entries are fetched, identified by their external id and a content hash, and embedded only when they are new or changed. An unchanged entry costs nothing.
Add your own API
You need a publicly reachable address that answers a GET request with a JSON list of objects. Addresses on private networks are rejected.
Name the base and say what it is for
The name and description are passed to the model — they are the only hint it has about when to reach for this base. Write them the way you would explain the base to a colleague.
Provide the address and authentication
Available methods: None, Bearer token, Header, and Query parameter. For a header or a query parameter you also supply the name the token is sent under. The token is stored encrypted and is never returned by the API — afterwards you only see that it is set.
Point at the list and map the fields
In Path to the list, give the location of the array of entries (for example data.items); leave it empty when the API returns a bare list. Then map five fields: Identifier, Title, Content, Entry URL, and Date.
The identifier must be stable — it is how we recognize the same entry across runs. Content goes into the embedding and into the model, so an entry without content is skipped. The entry URL is what lets an agent cite it as a source. A missing title is replaced by the first line of the content.
Test the connection
Test connection shows the raw source response side by side with the entries derived from it. The base cannot be saved without this step: the save button stays disabled until the preview returns at least one entry. Mapping blindly fails only during the first nightly sync.
Set the schedule and retention
Sync every: manually only, 1 h, 6 h, or 24 h. Delete entries older than: never, 30, 90, or 365 days — counted from when the entry was first seen.
Syncing and browsing the corpus
The base screen shows the state of the last run, how much of the plan quota is used, and the indexed corpus with a search box — so "why did the model not find this?" is answered by looking at the entries rather than by guessing. The button next to the name follows the state: Sync now, Try again, Resume, Cancel, or See skipped. While a run is in progress you see the page number and how many entries were created, updated, and skipped. The run also appears in the Queue as a knowledge sync job, but it does not take part in that queue's ordering, so it never delays writing.
Sync states:
| State | Meaning |
|---|---|
| Ready to use | The state of a live base — there is nothing to sync. |
| Never synced | The base exists, but no run has happened yet. |
| Sync running | A run is in progress. |
| Synced | The last run completed fully. |
| Synced, some entries skipped | The run succeeded, but some entries could not be mapped. |
| Sync failed | The last run ended with an error. |
| Automatic sync paused after 3 errors | Three consecutive failed runs. The scheduler stops querying the source. |
| Plan entry limit reached | Indexed entries stay; only new ones stopped. |
Adding a base is not enough — it must be attached
This is the most common misunderstanding. A base in the Library is only your catalog. Access comes from attaching it — in the list, a base with no attachments is labelled as invisible to the model.
Attach the base to a goal
Under Goal → Settings → Knowledge bases, select the bases that should be available to this goal's agents.
Enable the knowledge skill on the workflow node
Open the node in the workflow editor and give it the skill for querying knowledge bases. Only then does the node panel show the list of bases.
Leave the selection empty, or narrow it deliberately
A node with nothing selected inherits all of the goal's bases. Selecting specific bases narrows the node to those. A node can never reach a base the goal does not have — its list only ever offers bases attached to the goal.
If the last sync of an attached base failed, the panel warns next to it that the model will receive stale data.
Plan limits
| Plan | Knowledge bases |
|---|---|
| Free | Default bases, read-only. No own sources and no syncing. |
| Pro | Up to 10 sources and 20,000 entries. |
| Scale | Up to 30 sources and 200,000 entries. |
The entry limit is also the daily indexing cap: that is how many entries can go through embedding in a day.