Skip to content

Deployment

Radioso’s production contract is a multi-service deployment. The backend, frontend, and document worker are separate runtime units. Deployments that enable live dashboard updates add a dedicated realtime runtime and Redis or Valkey. The docs portal stays separate as well.

Production contract

  • independent image per service
  • independent Cloud Run service for backend, frontend, and worker
  • a dedicated realtime service and Redis or Valkey when live dashboard updates are enabled
  • Cloud SQL for application state
  • GCS for uploaded document source files
  • Cloud Tasks for request-driven worker dispatch
  • Secret Manager for runtime secrets

Required configuration categories

At minimum, a production deployment needs:

  • database connectivity
  • model provider credentials
  • session and token secrets
  • storage configuration for uploaded files
  • public base URLs for customer-facing flows such as public chat and Enterprise Edition modules
  • Operator MCP resource and issuer URLs when delegated Ray access is enabled

Relevant runtime flags are defined in .env.example and enforced by backend/src/app/config/env.ts.

Live dashboard updates

Live dashboard updates are optional. With REALTIME_MODE=disabled, dashboard queries keep their polling fallback and the realtime process and Redis are unnecessary.

For a small deployment, use standalone mode with one Redis or Valkey 7+ endpoint. Give the backend, workers, and realtime process the same mode, rollout, and Redis settings:

dotenv
REALTIME_MODE=standalone
REALTIME_REDIS_URL=redis://redis.internal:6379
REALTIME_ROLLOUT_MODE=default-on

Build the backend, then run the dedicated process with pnpm --dir backend run start:realtime. It exposes /health/live, /health/ready, and the internal /api/v1/events stream.

Set REALTIME_INTERNAL_URL on the frontend to the realtime service’s private base URL, for example http://realtime:8080. Browsers connect to the same-origin /backend/api/v1/events route; the frontend forwards that request to the internal stream without exposing the realtime service publicly.

The stream uses the signed dashboard session cookie and x-workspace-id. It is a browser transport endpoint, so personal/service API credentials, the TypeScript SDK, and MCP do not use it. Frames contain workspace identity and invalidation kinds rather than document or conversation content.

Radioso’s MCP server runs as a standalone HTTP process. Configure RADIOSO_BASE_URL and the standalone bind settings, then send the original MCP-audience agent credential to its /mcp endpoint. Hosted Terraform shares a generated RADIOSO_MCP_SIGNING_SECRET with the backend and configures Google’s two-hop forwarding suffix so each client keeps a separate pre-authentication budget. Manual deployments leave RADIOSO_TRUSTED_PROXY_HOPS=0 unless the rightmost proxy chain is controlled by the operator. The process may cache short-lived backend session tokens in memory; Postgres retains the conversation binding for the credential version, so restart and scale-to-zero cache misses resume that conversation. RADIOSO_MCP_REDIS_URL is optional for a shared cache, where the signing secret also encrypts cached tokens. See MCP server for the client flow.

Operator MCP

Operator MCP is the delegated OAuth resource for Ray’s workspace tools. It is a separate surface from an authored agent’s /mcp channel. Set the backend values together, then expose the same canonical resource through the frontend:

dotenv
OPERATOR_MCP_RESOURCE_URL=https://mcp.example.com/operator/mcp
OPERATOR_MCP_ISSUER_URL=https://app.example.com
OPERATOR_MCP_INTERNAL_SECRET=<at-least-32-character-secret>
OPERATOR_MCP_CREDENTIAL_EPOCH=1
RADIOSO_OPERATOR_MCP_PUBLIC_URL=https://mcp.example.com/operator/mcp

OPERATOR_MCP_RESOURCE_URL must be the exact /operator/mcp URL with no query or fragment. OPERATOR_MCP_ISSUER_URL is the public Radioso application origin that serves authorization and consent. The frontend value is read at request time and should match the resource the standalone MCP process serves. Operator MCP starts when those values, the internal secret, and the credential epoch are configured together. The GitHub Terraform workflow discovers the Cloud Run MCP URL unless MCP_PUBLIC_ORIGIN supplies a custom domain.

The credential epoch is a deployment-wide monotonic security value. Start at 1, and increase it when rotating the Operator MCP internal secret or invalidating all issued Operator MCP credentials. Before rolling replicas, load DATABASE_URL, OPERATOR_MCP_RESOURCE_URL, the new OPERATOR_MCP_INTERNAL_SECRET, and the new OPERATOR_MCP_CREDENTIAL_EPOCH; set OPERATOR_MCP_PREVIOUS_CREDENTIAL_EPOCH to the current persisted epoch; then run pnpm --dir backend run operator-mcp:rotate-credential-state once. Roll the same new epoch and secret to every backend and MCP instance before serving traffic. Instances with an older epoch or a different secret fingerprint fail readiness. Tokens and refresh lineages issued under an older epoch stop working, including after a database restore. Keep the epoch value at or above the last value used by the deployment.

The dashboard exposes operator:read, operator:probe, operator:propose, operator:write, and operator:act through the catalog in every workspace. operator:write applies only the exact reviewed operation the person confirmed: a draft-only change needs their confirmation in the conversation, and a live or irreversible change needs their signed-in approval on the review page. It does not replace the user’s workspace permissions. Verification is capped at six operations per credential each minute. The API access card reports disabled, misconfigured, or unavailable states without offering a personal-token workaround.

Enterprise Edition deployment

Enterprise Edition

Enterprise deployments use the same Cloud Run services, but the deploy workflow builds the images with radioso_edition: enterprise. In practice, that does two things:

  • the backend image installs @radioso/enterprise-backend-module
  • Terraform injects RADIOSO_APPLICATION_MODULES and edition flags into Cloud Run

Keep app_base_url_override set to the public frontend origin after the first service rollout, for example https://app.example.com. If you use the included Terraform workflow, set the GitHub environment variable APP_BASE_URL to that origin. If you deploy another way, set the backend runtime APP_BASE_URL to the same origin. Terraform also forwards this value as RADIOSO_WIDGET_ORIGIN. Password reset emails, email verification emails, user invitation emails, and the website embed snippet use these values when they build hosted links.

Once accounts start creating content, cap how much of it an account can index and how many answers it can generate — see Enterprise usage limits for the meters, the admin API, and the 429 responses callers see at a cap.

Required secrets

The Terraform workflow expects these GitHub environment secrets for each deployed environment:

  • OPENAI_API_KEY
  • SESSION_COOKIE_SECRET
  • WORKSPACE_TOKEN_SECRET
  • PUBLIC_CHAT_SESSION_SECRET
  • CONNECTOR_ENCRYPTION_KEY
  • RESEND_MAIL_API_KEY
  • POSTHOG_API_KEY

PUBLIC_CHAT_SESSION_SECRET signs public chat session tokens for anonymous chat links and website embeds. In Terraform this maps to public_chat_session_secret, which is stored in Secret Manager and injected into the backend as PUBLIC_CHAT_SESSION_SECRET.

WORKSPACE_TOKEN_SECRET derives the per-agent signing key a workspace admin can reveal so their website’s own backend can prove signed visitor identity for embedded chat. It is not a signer for personal or service API credentials.

Set MAIL_FROM_EMAIL to a verified sender address. MAIL_FROM_NAME is optional and defaults to Radioso. Without RESEND_MAIL_API_KEY, the backend writes outbound mail to the log instead of sending it, so password resets, email verification, and user invitations reach nobody.

The included staging and live Terraform environments enable PostHog error reporting with ERROR_SINKS=audit,posthog. Staging also sends product analytics to PostHog with PRODUCT_ANALYTICS_SINKS=audit,posthog. Set POSTHOG_API_KEY for each environment. POSTHOG_HOST defaults to the US ingestion host and can be overridden for another PostHog region or a self-hosted instance.

“Sign in with Google” on the Enterprise login page is optional and off until the environment carries GOOGLE_LOGIN_CLIENT_ID and GOOGLE_LOGIN_CLIENT_SECRET. The workflow stores both in Secret Manager and injects them into the backend. Register <APP_BASE_URL>/api/v1/ee/auth/google/callback as the redirect URI on the Google OAuth client; one client can list the callback for every region you deploy. A first Google sign-in joins the account that already holds the same email address and marks that address verified, so someone who registered with a password keeps the one account.

Deploy the frontend release carrying the callback route before you apply the credentials. The deploy workflow resolves TF_VAR_frontend_image from the frontend service already running and hands it to Terraform, which never advances it, so credentials applied first flip /api/v1/ee/auth/google/status to enabled and show the button while the browser still lands on a 404 on the way back from Google.

Alerting

Cloud Monitoring alerting is off until an environment asks for it. Set these on the GitHub environment and the next Terraform apply builds an uptime check on the backend /health route, a log-based error metric, and alert policies for Cloud Run 5xx rate, backend p95 latency, error-log volume, Cloud SQL saturation, Cloud Tasks backlog, and Cloud Scheduler failures.

ValueWhat it does
MONITORING_ENABLEDtrue creates the alerting resources for this environment
MONITORING_NOTIFICATION_EMAILSComma-separated addresses, one email channel each
MONITORING_EXTRA_NOTIFICATION_CHANNEL_IDSComma-separated IDs of channels you created in the Cloud Monitoring console

Give it at least one of the two destinations. When MONITORING_ENABLED is true and both are empty the workflow stops and names the value to set, because alerting that applies cleanly and delivers nowhere is worse than none.

Slack is a console channel rather than a Terraform one, because it holds an OAuth token that would otherwise sit in Terraform state. Create it once under Monitoring → Alerting → Edit notification channels, copy the projects/<project>/notificationChannels/<id> value, and put it in MONITORING_EXTRA_NOTIFICATION_CHANNEL_IDS. Every policy delivers to the union of the email channels and that list.

Ops event feed

Signups, chat turns, and errors reach an HTTP endpoint of your choosing — a Slack workflow trigger, n8n, or your own handler — when you add ops_webhook to a sink list:

ValueWhat it does
PRODUCT_ANALYTICS_SINKS / ERROR_SINKSAdd ops_webhook, for example audit,posthog,ops_webhook
OPS_EVENT_WEBHOOK_URLWhere each signed JSON delivery is posted
OPS_EVENT_WEBHOOK_SECRETEnvironment secret the delivery signature is computed with
OPS_EVENT_WEBHOOK_EVENTSComma-separated event names to forward; every product event goes out when unset

The backend refuses to start when a sink list names ops_webhook without both a URL and a secret, so the workflow checks for them first and stops the deploy rather than rolling out a revision that crash-loops. chat.started and chat.completed fire on every conversation, so name the events you want before pointing this at a channel someone reads.

Public chat rate limits

Public chat and website embed rate limits are operator tuning knobs, not workspace settings. When unset, the backend uses built-in defaults.

  • PUBLIC_CHAT_RATE_LIMIT_WINDOW_MS
  • PUBLIC_CHAT_SESSION_RATE_LIMIT_MAX_ATTEMPTS
  • PUBLIC_CHAT_GLOBAL_RATE_LIMIT_MAX_ATTEMPTS
  • PUBLIC_CHAT_SESSION_READ_RATE_LIMIT_MAX_ATTEMPTS

The session limit applies to a visitor session, and covers starting one and sending a message. The global limit applies across public chat traffic for the process. The read limit covers the same visitor session polling for history, tail, and live updates, which happens far more often than sending a message, so it carries a higher default. All three are process-wide: Radioso does not read a per-workspace override, so set them to values every workspace on this deployment can share.

Agent channel chat rate limits

REST agent chat and MCP ask calls each spend two durable limits before model work starts. One bucket follows the agent-channel credential, while the other covers all agent-channel credentials in the workspace. This keeps a single leaked credential contained and prevents a credential pool from consuming a workspace’s whole provider budget.

  • AGENT_CHANNEL_CHAT_RATE_LIMIT_WINDOW_MS defaults to 60 seconds.
  • AGENT_CHANNEL_CHAT_GRANT_RATE_LIMIT_MAX_ATTEMPTS defaults to 30 requests per credential and window.
  • AGENT_CHANNEL_CHAT_WORKSPACE_RATE_LIMIT_MAX_ATTEMPTS defaults to 300 requests per workspace and window.

Authenticated request limits

Authenticated assistant chat and the retrieval answer and search routes share one durable rate limit, because each request can trigger model or retrieval work. Tune it with EXPENSIVE_AUTHENTICATED_RATE_LIMIT_WINDOW_MS and EXPENSIVE_AUTHENTICATED_RATE_LIMIT_MAX_ATTEMPTS; the default is 60 requests per 60 seconds. Browser sessions share a bucket per account and workspace, and each workspace API token gets its own bucket within that same account and workspace, so one integration’s burst does not starve the dashboard.

Audience Pulse has its own budget: each account and workspace can make three explicit refresh attempts every 15 minutes. Opening the saved report costs nothing — only pressing refresh counts. Audience Pulse is available to cookie-authenticated dashboard sessions; workspace API tokens and bearer requests cannot reach it.

Ray turns draw on the same EXPENSIVE_AUTHENTICATED_RATE_LIMIT_* window, in a bucket per operator as well as per account and workspace. A Ray turn is a loop the operator starts and a model runs, so bucketing per operator keeps one busy session from refusing turns for a colleague in the same workspace.

Ray verification budget

Ray’s verification tools spend real model budget on every call: a test turn runs the agent for real, a case replay re-runs a recorded turn, and a suite run does that once per case. One Ray turn may spend six replayed turns, tunable with COPILOT_PROBE_BUDGET_PER_TURN, and a suite run is charged per case — asking for five cases spends five. When the budget runs out Ray answers with what it already measured and asks the operator to send another turn.

Each replayed turn also counts against the same answer meter a live customer turn does, so a suite run of five cases is charged as five answers. A suite that exhausts the workspace’s allowance part-way stops there and returns the quota refusal; the cases it already recorded are kept. On a Terraform-managed deployment the matching variable is copilot_probe_budget_per_turn, which passes the value to the API service where Ray turns run.

Ray conversation retention

Ray conversations hold operator questions, configuration detail, and excerpts of customer conversations, so they expire on a schedule. A conversation is deleted 90 days after its last activity, along with its messages and proposals. Set COPILOT_CONVERSATION_RETENTION_DAYS to change the window, or 0 to keep conversations indefinitely; the Terraform variable is copilot_conversation_retention_days.

The sweep runs in the document worker, so retention happens only where that process runs. Under the Cloud Run task runtime the trigger is a scheduled push to POST /internal/tasks/copilot-retention/sweep, which the bundled Terraform provisions as a daily Cloud Scheduler job. A deployment that runs neither the worker poll loop nor that schedule keeps Ray conversations forever, whatever the window says.

Slack inbound event retention

Every Slack webhook delivery is recorded by its event id, with no message text, so a retried delivery is answered once. The document worker deletes ids older than SLACK_INBOUND_EVENT_RETENTION_DAYS (default 7; 0 keeps them). Under the Cloud Run task runtime the trigger is a scheduled push to POST /internal/tasks/slack-inbound-event-retention/sweep, which the bundled Terraform provisions as a daily Cloud Scheduler job (override the cadence with slack_inbound_event_retention_schedule).

Agent bundle import recovery

An interrupted bundle import can leave the newly created agent behind before the request can compensate for a later failure. The document worker checks durable import jobs and removes agents from jobs that have stayed active longer than AGENT_BUNDLE_IMPORT_ORPHAN_AGE_MS. The default is 900000 (15 minutes); raise it only when an import can legitimately take longer in your deployment. Imports normally take seconds: once an import exceeds the threshold, cleanup fences it and the request returns a retryable conflict instead of claiming an agent that cleanup may delete. Terraform-managed deployments set the same value with agent_bundle_import_orphan_age_ms; Cloud Scheduler calls POST /internal/tasks/agent-bundle-imports/sweep every five minutes by default (override with agent_bundle_import_cleanup_schedule). Each cleanup writes an audit event and a agent_bundle_import_compensations_total metric, with ids but never bundle content.

External skill call timeout

A routine step that calls a skill’s MCP tool bounds that call with EXTERNAL_MCP_TOOL_CALL_TIMEOUT_MS — 30 seconds by default, up to 90 seconds. Connecting to the server and discovering its tools use a separate, fixed 10-second bound.

Reverse proxy client IPs

Express trust proxy stays off by default, and TRUST_PROXY_HOPS=0 is the secure setting for a deployment with no proxy in front of the backend. When the backend runs behind reverse proxies you control and rate limits must see the real client IP from X-Forwarded-For, set TRUST_PROXY_HOPS to the exact number of trusted hops. On Cloud Run behind the frontend API proxy that is typically 1 or 2, depending on the topology.

Runtime separation

The backend owns migrations, auth, retrieval, chat, and API routes. The worker owns queued document processing. When enabled, the realtime runtime owns long-lived dashboard streams. Keep those services independently deployable and independently scalable.

That separation matters because worker traffic behaves differently from chat traffic:

  • worker jobs are durable and retryable
  • chat is synchronous and user-visible
  • worker resource pressure should not degrade app responsiveness

In the hosted Cloud Run environments, the backend and frontend scale to zero when idle. The document worker also defaults to worker_min_instances = 0 and is invoked by Cloud Tasks when document work is queued. Raise worker_min_instances only when an environment needs a continuously warm polling fallback for queue recovery.

Worker dispatch

Document ingestion always writes a durable processing job to Postgres first; the dispatch driver only controls how workers learn that the job exists. Set WORKER_DISPATCH_DRIVER to one of:

  • noop — the default. Workers poll the job table. Right for local runs and single-host deployments.
  • cloud-tasks — the API enqueues a Cloud Task per job. Configure WORKER_TASKS_QUEUE_LOCATION, WORKER_TASKS_QUEUE_NAME, WORKER_TASKS_CRAWL_QUEUE_NAME, WORKER_TASKS_SERVICE_URL, WORKER_TASKS_INVOKER_SERVICE_ACCOUNT, and WORKER_TASK_AUTH_TOKEN. Cloud Tasks and Cloud Scheduler send that shared token in X-Radioso-Worker-Token alongside Google OIDC, so a proxy in front of a self-hosted task server must preserve the header. Generate a random value of at least 32 characters and give the backend, document worker, and crawler worker the same value.
  • amqp — a RabbitMQ-compatible AMQP 0-9-1 broker wakes the workers. Configure WORKER_AMQP_URL, WORKER_AMQP_QUEUE_NAME, WORKER_AMQP_CRAWL_QUEUE_NAME, and WORKER_AMQP_PREFETCH. Messages carry job ids and trace metadata only; Postgres remains the source of truth for job state, retries, and leases, and the worker polling loop stays active for recovery, so a lost broker message delays a job rather than losing it. Delayed retries follow the job table’s available_at value, not broker-delayed delivery.

Document processing and website crawling use separate queues in both push drivers, so a long-running crawl does not block document work.

A routine’s fire-and-forget actions — a contact request, a handoff notification, an approval request — go through their own durable outbox, routine_action_requests, with the same durable-first, dispatch-second shape: a worker drains it on a five-second poll by default. On Google Cloud, set ACTION_DISPATCH_TASK_QUEUE_NAME alongside the worker dispatch settings so a turn that enqueues an action pushes a Cloud Task immediately and reaches the worker in seconds rather than on the next poll. A scheduled recovery sweep drains anything a push loses or never sends.

Startup migrations

The backend runs SQL migrations before it binds the HTTP port. This is intentional: the API should not report healthy or serve traffic until the schema matches the running code.

Workers do not apply SQL migrations. They only check for pending migrations before starting document or crawl work. If a worker starts but the backend revision does not, treat the backend startup migration path as the first thing to inspect.

Startup migration metadata checks use separate lock and statement timeout budgets:

  • DB_MIGRATION_LOCK_TIMEOUT_MS
  • DB_MIGRATION_STATEMENT_TIMEOUT_MS

In practice, a blocked migration metadata check fails with an application log that names startup migrations as the failing phase, rather than looking only like a platform port-listen failure. A lock timeout usually means another database session is holding or waiting on the migration metadata table. The migration SQL body disables those local metadata timeouts so large index builds and backfills can still finish.

The backend owns migrations during startup; there is no separate pre-deploy migration job.

Agent revision rollout

Agent revision storage changes the order of an upgrade. Before applying migration 171, stop admission to replicas running the old application and stop old worker consumers. Wait for in-flight requests and claimed jobs to drain, then verify that no old replica remains. Apply the migration and inspect the published baseline pointers, draft state, and conversation bindings before starting revision-aware replicas. Do not run old and revision-aware writers or readers together: they do not share the same publication boundary.

Existing agents receive a migrated published baseline. New and imported agents remain private until an operator publishes them. A published revision is shared by every channel; a new conversation uses the current published revision, while an ongoing conversation continues with the revision it started on. There is no per-channel publication pointer.

Active routine state that refers to an older exact definition is retained only when its closure can be identified safely. Missing, ambiguous, or invalid routine pins are surfaced for operator action and fail closed; they are not silently reassigned to the current routine. Validation of this legacy-state retention is part of the migration readiness check, so do not declare the rollout complete from the presence of baseline rows alone.

Review the unbound conversations before starting revision-aware traffic:

sql
SELECT conversation_id, workspace_id, agent_id, routine_id, classification, created_at
FROM agent_revision_migration_classifications
ORDER BY created_at, conversation_id;

Each row names the original routine pin and why Radioso could not bind that conversation. Keep those conversations out of service until the pin has been investigated and the operator has a safe recovery path for that exact routine state.

Docs portal deployment

The docs portal is its own Next.js app under docs-portal/. It syncs the backend OpenAPI document during build and should be deployed as a separate public service instead of being mounted into the product frontend.

Rollout checklist

  1. Apply database migrations before or during backend rollout.
  2. Confirm backend can reach Cloud SQL and storage.
  3. Confirm worker runtime can claim or receive jobs.
  4. Confirm frontend points at the current backend service.
  5. If realtime is enabled, confirm Redis connectivity, realtime readiness, and the frontend’s exact /backend/api/v1/events proxy route.
  6. Confirm public URLs used for auth and embed flows are correct.
  7. Confirm docs portal build synced the current OpenAPI document.
i

If the app is live but document uploads never become searchable, treat that as a worker-path or dispatch-path incident first.

  • Self-hosting operations — backup, restore, and upgrade steps for deployments you run yourself, including the migration-startup incident playbook this page’s startup-migration behavior feeds into.
  • Enterprise usage limits — cap indexed storage, monthly content, and monthly answers per account.
  • Deployment topology — how the backend, frontend, and worker services relate at runtime.