Upgrade and recover a deployment
Preview a workspace upgrade, deploy with a recovery plan, verify the result, and choose the correct rollback path.
7 min read
For a workspace deployment, tale update changes the CLI and workspace files; tale deploy changes the running services. Choose the target version and recovery point before either step. A blue-green rollout overlaps application replicas, but snapshots, drain periods, and stateful-service replacements can interrupt work.
Managed deployments use pinned source revisions and prepared bundles instead of this workspace procedure. Follow Managed deployments for that workflow, or Configuration releases when only client content changes.
Prepare the upgrade
- Run
tale statusin the intended workspace. Record the running version, workspace version, and current deployment state. - Read the target release notes for compatibility, required configuration, and known limitations. A pre-0.5 instance needs the separate cutover below.
- Confirm a restorable off-host backup, the matching keys, and coverage of external databases and buckets. Backups and restore defines the recovery set.
- Allow capacity for old and new application replicas at the same time. Agree on a maintenance window when snapshots, sandbox replacement, or stateful updates can interrupt required work.
- Preview the selected update and deployment. Inspect warnings rather than treating a successful preview as proof that a live migration will succeed.
Select the version
Workspace instance commands try to align the CLI with the version recorded in tale.json. If downloading that version fails, the CLI warns and continues with the current binary. Resolve an unexpected mismatch before making deployment changes.
Without a version argument, tale update selects the newest release in the workspace's current major.minor line. Moving to another line is explicit:
tale update --dry-run
tale update --version <target-version> --dry-run
tale update --version <target-version>The update replaces the CLI and synchronizes workspace templates, leaving running containers alone. If file synchronization fails, it attempts to return the binary to the workspace's prior version. Review the resulting files and output before deploying.
Preview and deploy
tale deploy --dry-run
tale deploy
tale statusA version-changing deployment or host-config override takes a local snapshot before mutation unless --skip-backup is supplied. This snapshot is additional protection, not an off-host recovery plan.
| Service group | Ordinary deployment | When to plan extra interruption |
|---|---|---|
platform, backend-api, backend-worker | Roll together as the new application colour. | Old and new replicas overlap, and draining can refuse new turns. |
sandbox, sandbox-egress, sandbox-llm-gateway | Replace in place after draining relevant work. | These are shared execution dependencies, not a second blue-green application group. |
db, object-store, proxy | Keep the running services; the CLI reports skipped updates. | Add --stop when these services need replacement. |
tale deploy --stopThe role replica variables TALE_PLATFORM_REPLICAS, TALE_BACKEND_API_REPLICAS, and TALE_BACKEND_WORKER_REPLICAS accept 1–16. Increase the role that measurements show is constrained; adding workers does not solve an unavailable database or provider quota.
Understand the handover
The CLI starts the idle colour and waits for its replicas to pass health checks before completing the handover. Both versions can serve during the overlap, so releases must remain compatible with the previous application version while migrations run.
The old API is drained before removal: new chat turns can receive a drain refusal while existing turns get time to finish. The chat drain waits up to three minutes; the web drain uses DRAIN_TIMEOUT, which defaults to 30 seconds. The web health route stays healthy while its alias is still shared. Disconnecting the old containers from serving networks removes them from DNS and can sever remaining connections, so it follows the drains.
Browser tabs opened before the handover still run the previous version. The first time such a tab needs a part of the application that the new version replaced, such as a document preview, it reloads once and continues on the new version. If that part still cannot load after the reload, for example because the tab reached the old colour again, the tab does not reload a second time. It shows A new version is available with a Reload action instead. A tab that cannot reach Tale at all does not reload; it shows its connection notice until Tale answers again.
If the new group does not become healthy within HEALTH_CHECK_TIMEOUT, the deploy does not complete the flip. Inspect the recorded deployment state and logs before retrying. An interrupted rollout can leave both groups or pending handover state; use the CLI's recovery output rather than deleting containers or state files by hand.
Check migrations and the user outcome
At boot, the backend applies its numbered migrations in file-name order under a session advisory lock. SQL migrations change the schema; TypeScript data migrations update existing rows by the application's own rules. The app_migrations table records both kinds by file name, so each runs once per database. Other replicas wait for that migration path. A migration error prevents the new backend from starting normally; inspect its error and the database before retrying. Forward-only migrations are not undone by changing an image tag.
tale migrate refreshes built-in organization defaults; it is not a command for rolling database migrations backward. Review whether local configuration was meant to be replaced before using host-config override options.
After deployment, verify the public certificate and sign-in, open an existing project or conversation, download a known file, and run a controlled check of the knowledge and automation paths you use. Check worker progress, store health, and the final deployed version. Keep the pre-upgrade recovery set until the deployment has met your acceptance criteria.
Choose a rollback path
| Situation | Recovery path |
|---|---|
Return to the recorded previous version in the same major.minor line | tale rollback checks that boundary and asks for confirmation before redeploying. Review that release's compatibility notes as well. |
| Return across a minor or major boundary | Restore the coordinated pre-upgrade data and deploy its matching version. tale rollback refuses this image-only downgrade. |
| Target version or data compatibility is unknown | Resolve the version and backup provenance before starting an older binary. |
tale rollback--yes skips its confirmation for an already approved unattended operation. The CLI's same-line check is a version guard, not an independent proof that every external integration or locally customized configuration is compatible. Never assume that downgrading is safe merely because an old migration list is a prefix of the new one.
0.4 → 0.5: a separate installation
The 0.5 application store replaced the earlier Convex database with Postgres. There is no in-place importer between those stores. Keep the old instance and its backups intact while preparing a fresh deployment in a separate workspace and data set.
Recreate organizations and users, review and transfer compatible configuration, and reimport required documents. Files left in an external bucket do not automatically acquire references in the new application database. Accept the replacement environment before decommissioning the old one.
The CLI refuses the unsupported cutover by default. Its expert --accept-data-loss override is not a migration tool and must not be used to preserve old application data. Historical volumes or databases can remain after earlier upgrades; their presence alone is not a reason to delete them during this procedure.
0.3 → 0.4: the OpenAI-compatible API was removed
From 0.2.10 through 0.3, Tale served an OpenAI-compatible layer under /api/v1: POST /api/v1/chat/completions and POST /api/v1/images/generations in the OpenAI request and response shapes, and an OpenAI-shaped GET /api/v1/models. Its model field could name an agent. The 0.4 rebuild removed this layer, and no later release restores it. Callers of these routes, including OpenAI SDKs pointed at the instance, stop working, because 0.4 serves none of the three routes. A current release answers chat completions and image generations with 404 NOT_FOUND, or with 400 ORG_SLUG_REQUIRED when the key holder belongs to several organizations and the request carries no X-Organization-Slug, which an OpenAI SDK does not send by default. Its GET /api/v1/models is Tale's own listing of models and agent harnesses, which an OpenAI client cannot read.
Find those callers before the upgrade and plan their replacement. Scripted questions move to the asynchronous REST chat API, which answers as the workspace assistant rather than as a bare model. Editor integrations that wanted Tale's knowledge use the MCP endpoint with the editor's own model, and work that must run on the organization's models goes to a project agent on a task. Use Tale from your editor or a script describes each path.