Skip to content

Self-hosting operations

Self-hosted Radioso has two durable data surfaces:

  • PostgreSQL, including accounts, workspaces, settings, conversations, audit events, documents, chunks, vectors, and worker queues
  • uploaded source-file storage, either the local document-storage volume or the configured object-storage bucket

Back up both surfaces together. A database backup without source files can leave document records whose original uploads cannot be reprocessed. Source files without the database are not enough to restore accounts, settings, embeddings, or chat history.

Redis or Valkey used for live dashboard updates is transient. Restore PostgreSQL and source files, then let the realtime runtime rebuild subscriptions and admission leases as browsers reconnect.

Minimum operating baseline

For a small private deployment, the minimum baseline is:

  • run the backend, frontend, document worker, crawler worker, PostgreSQL, and document storage as separate services or containers
  • put the frontend and backend behind HTTPS
  • keep the default PostgreSQL host publication on 127.0.0.1; use the private container network or a private database network for service access
  • give the backend and HTTP task workers the same WORKER_TASK_AUTH_TOKEN, and preserve X-Radioso-Worker-Token through any reverse proxy
  • keep .env and secret values outside source control
  • back up PostgreSQL and document storage every day
  • test restore before relying on the backup
  • deploy a tagged release and pin the image to it, rather than tracking a branch
  • apply upgrades manually after reading the migration list

Live dashboard updates are optional. An enabled deployment also runs Redis or Valkey 7+ and the dedicated realtime process described below.

Backup

For Docker Compose deployments, back up:

  • the PostgreSQL database with pg_dump or a physical volume snapshot
  • the radioso_document_storage volume when DOCUMENT_STORAGE_DRIVER=local
  • the deployed .env or equivalent secret inventory, stored in a secure password manager or secrets system

For cloud deployments, back up:

  • Cloud SQL or the managed PostgreSQL database
  • the GCS bucket or object-storage bucket used by DOCUMENT_STORAGE_BUCKET
  • Secret Manager values or the external secret inventory

Keep database and source-file backups from the same time window. That makes document reprocessing predictable after restore.

Restore

Stop writers

Stop the backend, document worker, crawler worker, and realtime process before restoring. This prevents new writes while the database and file storage are out of sync and closes active dashboard streams.

Restore PostgreSQL

Restore the database first. Confirm the schema_migrations table exists and matches the deployed code version.

Restore source files

Restore the local document-storage volume or object-storage bucket. Keep object paths unchanged.

Start the backend

Start the backend and confirm /health returns success. The backend owns SQL migrations, so it should be the first runtime to touch the restored database.

Start workers

Start the document worker and crawler worker. Watch for pending or failed processing jobs before opening the deployment to users.

Start realtime and the frontend

If realtime is enabled, start Redis or Valkey and the realtime process, then wait for /health/ready. Start the frontend after its internal realtime upstream is reachable.

Upgrade

Use this order for self-hosted upgrades:

  1. Read the release notes for the version you are moving to. They list every SQL migration the release adds.
  2. Back up PostgreSQL and document storage.
  3. Build or pull the new backend, frontend, worker, and optional realtime images.
  4. Start the backend and let migrations run.
  5. Start the document worker and crawler worker.
  6. If enabled, start Redis or Valkey and the realtime process and confirm /health/ready succeeds.
  7. Start or refresh the frontend.
  8. Upload a small document and confirm it becomes searchable.
  9. Ask one grounded question and verify citations or retrieval trace data.
  10. curl <backend_url>/health and confirm version matches the release you meant to deploy.

/health reports the build it is running:

json
{ "status": "ok", "version": "1.4.0", "commit": "5434e0eb6f2c1d..." }

version is the release the image was built from. A build made outside a release reads as <last release>+<short commit>, and a build with no deploy stamp at all reads as development — so a service can never claim a release it does not carry. When you run more than one stack, curl each of them: this is the fastest way to see that one is behind another.

Do not auto-pull main into production without a backup. A failed application upgrade is usually reversible; a partially applied schema or mismatched storage restore is harder to recover.

The backend runs SQL migrations during startup before it opens the HTTP port. The document worker and crawler worker only check for pending migrations. Start the backend first, then start workers after the backend has completed migrations and /health responds.

A rolling deploy or a multi-replica start can bring up two backend instances at once. Before touching schema_migrations, each instance takes a session-level PostgreSQL advisory lock, so exactly one instance applies the pending migrations while the rest wait their turn.

Migration startup metadata checks have their own timeout controls:

  • DB_MIGRATION_LOCK_TIMEOUT_MS
  • DB_MIGRATION_STATEMENT_TIMEOUT_MS

Keep these values shorter than your platform startup-probe window. That way a blocked migration metadata check appears as a clear application log instead of a silent port-listen timeout. A waiting instance that exceeds DB_MIGRATION_LOCK_TIMEOUT_MS exits with a log line naming the migration lock instead of hanging past the probe; the platform’s restart then finds the migrations already applied by the instance that held the lock and starts cleanly. Large migration SQL bodies, such as index builds or backfills, are not capped by these local metadata timeouts.

Live dashboard updates

The default REALTIME_MODE=disabled needs no Redis service and leaves dashboard freshness to polling. To enable live updates for a small deployment, use the same configuration on the backend, workers, and realtime process:

dotenv
REALTIME_MODE=standalone
REALTIME_REDIS_URL=redis://redis:6379
REALTIME_ROLLOUT_MODE=default-on

Build the backend and run pnpm --dir backend run start:realtime. Configure the frontend with the realtime process’s private base URL:

dotenv
REALTIME_INTERNAL_URL=http://realtime:8080

Keep the realtime process on the private service network. The browser opens GET /backend/api/v1/events on the frontend, and the frontend forwards it to the realtime process’s /api/v1/events route. Preserve the request’s cookie and x-workspace-id headers when placing another proxy in front of the frontend.

This stream accepts the dashboard session cookie rather than personal or service API credentials. It is used by the dashboard itself and is not an SDK or MCP event subscription API.

For rollback, set REALTIME_MODE=disabled and REALTIME_ROLLOUT_MODE=disabled across the application runtimes, then stop the realtime process and Redis. Visible dashboard queries continue through their polling fallback.

Migration startup incidents

If a new backend revision fails before /health is reachable, check the backend logs first. A migration metadata lock or statement timeout should name startup migrations as the failing phase.

Look for database migration lock held by another instance; waiting — a normal line during a rolling deploy, logged once by the instance that lost the race for the session-level migration lock. If that instance’s log ends with Timed out waiting for the database migration lock held by another instance (DB_MIGRATION_LOCK_TIMEOUT_MS), the instance holding the lock took longer than DB_MIGRATION_LOCK_TIMEOUT_MS to finish, or never released it. Check that instance’s own logs for the migration it was running when it stalled.

If the worker starts but the backend does not, that is a useful signal. Workers use a read-only pending-migration check, while the backend is the runtime that applies SQL migrations.

To inspect blocking sessions in PostgreSQL, use a privileged database console and check locks around the migration metadata table and the session-level migration advisory lock:

sql
SELECT a.pid, a.state, a.wait_event_type, l.mode, l.granted, left(a.query, 140) AS query
FROM pg_locks l
JOIN pg_class c ON c.oid = l.relation
JOIN pg_stat_activity a ON a.pid = l.pid
WHERE c.relname = 'schema_migrations';
 
SELECT a.pid, a.state, l.granted, left(a.query, 140) AS query
FROM pg_locks l
JOIN pg_stat_activity a ON a.pid = l.pid
WHERE l.locktype = 'advisory';
 
SELECT pid, pg_blocking_pids(pid) AS blocked_by, state, left(query, 140) AS query
FROM pg_stat_activity
WHERE cardinality(pg_blocking_pids(pid)) > 0;

If a stale session is clearly blocking a rollout, terminate only that backend session:

sql
SELECT pg_terminate_backend(<pid>);

Use this carefully. Confirm the session is stale or belongs to the failed rollout before terminating it. Restarting the database is the broader fallback when session-level recovery is not clear.

Worker incidents

If uploads succeed but new documents do not become searchable, treat it as a worker-path incident first.

Check in this order:

  1. The backend can accept uploads and write document records.
  2. The document worker is running and can connect to PostgreSQL.
  3. Source-file storage is reachable from the worker.
  4. The embedding provider credentials and model settings are valid.
  5. The queue is draining rather than growing.
  6. Failed jobs have enough logs to identify parser, storage, embedding, or timeout failures.

For local polling deployments, WORKER_DISPATCH_DRIVER=noop means the worker claims durable jobs from PostgreSQL. For queue-backed deployments, also check Cloud Tasks or AMQP delivery and retry configuration.

Secrets

Rotate these values deliberately and keep historical impact in mind:

  • SESSION_COOKIE_SECRET affects browser sessions.
  • WORKSPACE_TOKEN_SECRET derives the per-agent signing keys used for signed visitor identity.
  • PUBLIC_CHAT_SESSION_SECRET affects anonymous public-chat sessions and website embeds.
  • CONNECTOR_ENCRYPTION_KEY protects both connector secrets and per-workspace LLM provider API keys at rest. The value must be 32 random bytes, base64-encoded — generate one with openssl rand -base64 32. The bootstrap command generates a key automatically when .env does not already have one. Workspaces that have stored an API key in the UI cannot read it back after the key is changed; operators must coordinate a rotation by clearing or re-entering each affected provider credential.
  • RADIOSO_EDGE_PROOF_SECRET lets the frontend’s server-side proxy vouch for a public chat or embed visitor’s IP address and geo headers when it forwards a request to the backend. At least 32 characters, shared verbatim by both services. Leave it unset and the proof path never runs: the backend still serves every request, it just records that visitor’s IP, country, region, and city as unknown instead of trusting an unsigned header. docker-compose.yml and docker-compose.dev.yml already set a matching dev value on both services, so the local stack exercises the signed path with no extra setup.
  • provider API keys affect chat, rewrite, rerank, and embedding calls.

Changing encryption or token secrets without a migration or rotation plan can invalidate existing sessions, tokens, or saved connector credentials.

Visitor request facts

Two more settings shape what a conversation records about the request that started it.

RADIOSO_TRUSTED_PROXY_HOPS turns a forwarded-for chain into one client address: it’s the count of trusted hops to strip off the right end of the chain before trusting what’s left. The same value already governs the MCP server’s and the agent channel’s rate-limit source budgets. Set it to 1 behind a single reverse proxy, such as the bundled frontend proxy in front of the backend, or 2 when a load balancer sits in front of that proxy in turn — each hop appends one more address, so counting wrong reads the wrong entry as the client.

Country, region, and city need no setting of their own. Radioso reads whichever geo header the deployment’s load balancer, CDN, or reverse proxy already stamps on the request: Cloudflare’s CF-IPCountry, Vercel’s x-vercel-ip-country family, and Google App Engine’s X-AppEngine-Country all work unmodified. Any other proxy can rename its own header to x-client-region or x-client-city to be picked up the same way.

Health checks

At minimum, monitor:

  • backend /health
  • frontend reachability
  • worker process liveness
  • realtime /health/live and /health/ready when enabled
  • stream connection counts, rejected admissions, reconnect rates, and Redis transport loss when realtime is enabled
  • PostgreSQL connection errors
  • document-processing failures
  • embedding provider failures
  • disk or bucket storage errors
  • HTTP 429 and 5xx rates for public chat and website embeds
i

The key point is that self-hosting is mostly database, storage, worker, and provider-key operations. The app itself is straightforward once those surfaces are backed up and monitored.

  • Deployment — the production contract, required secrets, and rollout checklist this runbook assumes.
  • Enterprise usage limits — cap indexed storage, monthly content, and monthly answers per account on Enterprise deployments.