Skip to main content

Choose and move data stores

Separate deployment defaults from organization storage, configure knowledge and file connections, and plan existing-data migration.

9 min read

Choose storage for three kinds of data: application records, searchable knowledge, and original files. Moving one does not move the others. Storage placement also does not determine where a model provider or connector processes a request; include those destinations in your residency assessment.

Choose the scope of the change

StoreDeployment-wide settingOrganization-specific setting
Application database: users, chats, runs, and audit dataDATABASE_URLNo separate application database setting on this page.
Knowledge database: extracted text, embeddings, search indexes, and crawled contentKNOWLEDGE_DATABASE_URLSettings > Data residency > Knowledge database
Original files: documents, attachments, audio, and generated mediaOBJECT_STORE_*Settings > Data residency > Object storage

The packaged stack puts tale_app and tale_knowledge in one Postgres service, while keeping them separate databases. Other layouts can use separate services. The environment values generated by the deployment select the defaults; a bare application process does not invent working object-store credentials.

Organization settings require Admin or Owner permissions. Without an organization-specific connection, Tale uses the deployment default and keeps organization data scoped. An invalid configured knowledge connection fails rather than silently using a different database.

Prepare an external database

Provision the database and credentials before changing Tale. The application database needs a role with permission to apply its schema migrations. The knowledge database needs vector installed; pg_search enables the BM25 part of hybrid search. A target with only pgvector provides vector search without that keyword leg. Tale prepares knowledge schemas and tables but does not install extensions on a database you own.

Use a direct or session-compatible Postgres connection. Transaction-mode pooling is incompatible with session locks, LISTEN, and prepared statements used by the stack. For sslmode=verify-ca or verify-full, supply the required PEM roots through POSTGRES_CA_FILE and mount the file in the backend roles. Check connectivity and permissions from the deployment network.

Before a cutover, choose how existing application or knowledge data will reach the destination, and take coordinated backups. Saving a database connection only changes where subsequent operations go; it does not copy old rows or reindex old documents automatically.

Change deployment defaults

Update .env, then deploy or recreate the affected backend services. docker compose restart keeps the old environment. If you are moving both databases out of the bundled service, preserve its old volume until the new stores and recovery plan are accepted.

For the default object store, the backend reconciles default/object-storage/connection.json and its secret sidecar from the environment at startup. The boot result distinguishes seeded, reconciled, skipped when credentials are absent, and ignored for an operator-managed file. "managedBy": "operator" makes you responsible for that file instead of environment reconciliation.

Changing the default bucket or endpoint does not copy existing blobs. Copy them with storage tooling, retaining keys and required metadata, and coordinate the switch before removing the old store. An external database or bucket is outside the CLI volume snapshot's data coverage; update Backups and restore procedures as part of the change.

Connect an organization's knowledge database

  1. Open Settings > Data residency in the target organization and configure Knowledge database with host, port, database, user, SSL mode, and password.
  2. Use Test connection. It first checks the fields the way Save does and names a missing host, database, or user under its field instead of running. Check both connectivity and extension availability; a successful connection does not prove the old corpus has migrated.
  3. Save only when the destination and existing-data plan are ready. Subsequent requests use the selected connection without a container restart.
  4. Index a controlled document, search for a known phrase, and verify that pre-existing content required by your users remains available.

The files live under $TALE_CONFIG_DIR/<orgSlug>/knowledge/: connection.json, connection.secrets.json, and embedding.json. Secrets use SOPS when an age key is configured. Removing the connection returns routing to the deployment default; it leaves data in the external database and makes that data inaccessible through the removed connection.

Match the embedding model to the corpus

In Embedding model, choose the provider and stored credential, then the model. Model lists the embedding models the provider's catalog carries and fills the vector width the catalog states. A provider that offers no embedding model at all, such as Anthropic, cannot be chosen: the row reads Cannot embed, and you pick another provider. Where Tale knows no vector width for the provider, type the model tag (for Azure OpenAI, your deployment name) and the vector width that model produces. That applies to a shipped provider whose catalog lists no embedding model, to a provider your organization defined itself whose model listing carries none, and to Azure OpenAI, whose deployments carry your own names. The provider's embedding declaration decides between the two, and saving refuses a provider declared unable to embed however the request arrives. If Tale cannot check the declarations, the section says so and offers Try again; until the check answers, no model can be chosen. An optional base URL selects an OpenAI-compatible endpoint. Without a configured embedding model, knowledge indexing and search cannot operate normally.

Vector width is pinned per database on first use. Organizations sharing a database must use that width; a different width needs a separate compatible database. Changing the model can also make existing vectors incompatible even when the width stays the same. Plan reindexing with the chosen model instead of mixing embeddings blindly.

embedding.json can set minSimilarity, the assistant search's vector-leg floor; its default is 0.45. The settings form preserves this file value but does not expose a field for it. Tune it against representative queries. REST knowledge search applies a floor only when its request supplies one; this is not a universal score threshold for all searches.

Pace requests to a self-hosted embedding server

Two more optional embedding.json settings describe how much work the embedding server can take. Set them when you run the server yourself, for example a model server on your own hardware that computes one request at a time and queues the rest. Like minSimilarity, they exist only in the file: the settings form keeps them when you save, and the CLI declares them in the knowledge-embedding resource.

  • maxConcurrentRequests (1 to 64, default 3) is how many embedding requests to this model Tale keeps in flight at once for the organization. Document indexing, the indexing of incoming email, website scans and searches share this limit. Further requests wait in arrival order, except that a search query goes ahead of waiting indexing batches. Each Tale process counts separately, so the API and every worker replica can each reach the limit. A lower value takes effect at once; a higher one once the requests started under the old value have finished. On a server that computes one request at a time, a higher value adds no load; each request only waits longer.
  • minTokensPerSecond (any positive number) is the slowest rate at which the server computes embeddings for this model under its usual load. Measure it while other work runs on the same hardware, such as a chat model, but leave out the time a request waits behind other requests. Tale adds that waiting time itself.
json
{
  "providerSlug": "local-embedding",
  "model": "example-embedding",
  "dimensions": 1024,
  "baseUrl": "https://embeddings.example.internal/v1",
  "maxConcurrentRequests": 2,
  "minTokensPerSecond": 800
}

Each embedding request has a ceiling: 15 minutes for indexing, and 5 minutes for a search query, which a chat answer waits on. Without minTokensPerSecond, Tale cannot tell how long the server's queue may take, so a request may use its whole ceiling. With it, Tale gives each request time for the work that can be ahead of it or beside it on the server, plus its own. That work is the request's own tokens, estimated generously from its characters, plus the other requests this Tale process may have in flight and maxConcurrentRequests more from other clients. Tale counts each of those requests as at least a full batch of 64 texts with 1,024 tokens each, divides the total by minTokensPerSecond, and adds 50%, but never allows less than 60 seconds or more than the ceiling. With the example above, a full batch of ordinary text gets about eight minutes, and a search query gets its five-minute ceiling. A search query also ends after five minutes in total, including any wait for a free slot. If more clients share the server than one more Tale process with the same limit, state a lower rate.

A request that runs out of time is not sent again at once: the server had it the whole time, and a repeat would only lengthen its queue. A refused connection, a rate limit or a server error is retried after a pause that grows with each attempt, or after the pause a busy server asks for with Retry-After, up to one minute. A server that asks for a longer pause is left alone. When one batch of a request made of several batches fails, for example on a long web page, Tale cancels the other batches, whether they are running or still waiting, so the server stops working on them.

Indexing a document gets at most 15 minutes per attempt. When a large document or a long queue needs more, the attempt stops at that limit and cancels its open request; a request that runs out of time ends the attempt too. The next attempt starts after a pause that grows with each attempt and continues after the chunks already stored. If a document is still unfinished after six attempts, it shows as failed; Retry indexing continues from the stored chunks.

Connect an organization's bucket

  1. Provision an S3-compatible bucket and the required object permissions. Configure CORS for the actual browser origins and the needed GET, PUT, and HEAD methods.
  2. In Object storage, enter region, endpoint when needed, bucket, optional key prefix, and credentials. Use path-style addressing when your store requires it.
  3. Run Test connection, then save. A missing region or bucket, or a value past a field's length limit, is named under its field before anything is sent. The server test writes, reads, and deletes a test object; it does not test browser CORS.
  4. Upload and download a controlled file in the browser before relying on the new connection.

New uploads use the organization's bucket. Earlier default-store files can remain readable through mixed references, so connecting the bucket does not by itself meet a requirement to relocate history. Configuration lives under $TALE_CONFIG_DIR/<orgSlug>/object-storage/connection.json and connection.secrets.json.

Removing this connection sends new uploads to the default store. Existing objects remain in the organization's bucket but require the connection to be restored before Tale can read them again.

Move existing files deliberately

With the organization bucket saved, use Move existing files in the same section. Review any preview, confirm the move, and follow its progress. Keep the connection stable until it finishes.

The backfill walks that organization's referenced documents and history, uploaded files, synthesized audio, and video transcripts. It copies each blob with its content type, checks the destination size, then deletes the source copy. It preserves object keys and can resume a previously verified copy without duplicating the move. This is a move, not an additional backup or a cryptographic content audit.

The destination must differ from the default store. If the run fails, inspect its last error and progress before starting another; preserve both stores until the final result and representative old-file downloads are verified. See Secrets with SOPS for protecting the connection sidecars.

© 2026 Tale by Ruler GmbH — ISO 27001 & SOC 2 certified.

Tale is MIT licensed — free to use, modify, and distribute.