Ray
Ray, your operator copilot, helps you investigate an agent without leaving the dashboard. Ask why a conversation received a particular answer, why retrieval did not run, which trace stage produced a refusal, which documents failed to process, or which skills an agent can reach. Ray reads the workspace data available to your operator session and streams its findings as it works.
Ray checks your current workspace permissions again when it is about to read a protected record or draft, run, or apply a change. If access changes while a turn is in progress, that step stops without exposing its result. A pending proposal never carries permission to apply itself: the operator applying it needs the required permission at that moment.
When you ask about configuration, Ray can discover the workspace’s agents and read the selected agent’s settings, authored directives, and built-in answer directives. It can also list that agent’s routines and inspect a routine’s authored definition, so questions about overlapping directives or a routine that never triggers can start from the agent name rather than an internal id. When the answer is a change, Ray drafts it for you to review.
Ray can search the indexed evidence in your documents, then inspect one document’s chunks when the question is why a passage was missed. One read returns up to ten complete chunks in order, including their boundaries, metadata, indexed search text, and whether each chunk has an embedding for the workspace’s active model. Ray follows the next chunk index when it needs another page, so chunk text stays intact instead of being clipped to fit one large result.
Between those two reads sits the question they cannot answer on their own: what
does retrieval actually return for this query, for this agent. Ray runs that
search with the agent’s own source scope, retrieval settings, and answering
instruction, and reports the chunks it would ground on with their scores. The
result always names the agent it measured, so a probe on one agent is never
mistaken for another agent’s behavior or for the workspace defaults. If the
agent answers without retrieval at all, the probe says so — otherwise an empty
result reads as a knowledge gap when the real cause is a setting. Scoping needs
workspace.agents.read alongside workspace.retrieval.query. Each probe counts
against your per-operator limit for calls that spend model budget.
With workspace.documents.manage, Ray can reprocess one document or the
existing documents in one source, and it can recrawl an existing website
source. These are maintenance acts, so they queue immediately and appear in the
turn’s activity instead of becoming proposal cards. Recrawl uses the source’s
stored URL, page limit, and policy. Creating a source, supplying a different
URL, and reprocessing the whole workspace stay in the Knowledge Base and
Settings surfaces.
For a skill, Ray sees every setting’s name. It also sees the value of a setting only when that skill’s capability opts the setting in as safe to show — a retrieve skill’s source scope, metadata rules, vector top K, retrieval strategy, and instruction text; an email skill’s draft-or-send mode. Every other setting stays name-only, including a notify skill’s recipient list and webhook URL, so Ray can tell you that a notify skill has no delivery destination configured without ever reading the address or URL you put there.
Ray also reads the workspace’s context variable definitions and which of them an agent has enabled — the source (pushed, browser, or a resolver skill) and how the variable surfaces to the model. It never reads a variable’s value for a particular visitor session or customer, because that value belongs to a conversation Ray was not asked to inspect.
Open the panel
Ray is available on every dashboard view. Choose Ray in the top bar, or press Cmd/Ctrl+J. Ray stays available while you move between Inbox, Agents, Knowledge Base, Audience Pulse, and Quality.
The panel puts the conversation first:
- The thread shows your question, Ray’s streamed answer, the activity behind the answer, and any entities Ray read during the turn.
- Looking at sits above the thread and names the dashboard view and agent Ray uses as context. Expand it to see the customer conversation and the other entities currently on screen.
- The conversation menu in the panel header holds your saved Ray threads. Open it to start a New conversation, switch to a Recent conversation, or Delete the current one.
The full-page Ray view shows the same thread with the conversation list and context rail alongside it, which gives more room for a longer investigation.
Connect an external engine to Ray
An MCP-compatible engine can use Ray’s operator tools through Settings → API access → Radioso MCP. The connection uses the separate /operator/mcp resource and OAuth consent, so the external client receives authority for one user, one workspace, one client identity, and one revocable grant. Setup selection is not connection state; the API access card lists a grant only after the client completes consent.
The MCP catalog has this boundary:
operator:read— read workspace and agent state that the user can already access, including workspace settings andproposal_detail, which returns a dashboard-safe proposal preview and tells the client whether it is a reviewed operation.operator:probe— run bounded diagnostics, including a retrieval probe against an agent’s configured source scope and retrieval settings.operator:propose— prepare a reviewable routine, retrieval, publication, document, agent-settings, ingestion-settings, or directive operation without applying it.prepare_directivesupportscreate,edit,set_enabled, andremove; supplied structured fields stay verbatim, while coaching and the advisory coherence check run only during preparation. Removal is irreversible and shows references to the directive, so disabling is the reversible alternative. Directive changes remain an agent draft until publication. A directive create is fenced only on its agent existing, so an unrelated agent change made after preparing it does not invalidate it. Reviewed directive edits, enablement changes, and removals use the directive version and become stale after another change to that directive executes. Preparation and execution require the target’s manage permission.prepare_agent_settingsapplies all named settings for one agent together after confirmation; its review identifiescustomInstructionas a draft that needs publication, while other fields are live at execution.prepare_ingestion_settingsaffects documents processed after execution; useprepare_document_reprocessfor documents that are already indexed.prepare_document_importaccepts 100 documents with bodies up to 20,000 characters, requires stable external document ids, limits the complete JSON input to 2,000,000 UTF-8 bytes, and checks stored-document capacity before it writes.prepare_document_removalrecords the exact documents and their versions.prepare_document_reprocessreviews one tagged selector and its eligible and skipped counts.operator:write— execute, inspect the outcome of, or cancel one reviewed operation bound to the same grant and client. These tools reach only an operation aprepare_*tool created, and each checks the caller’s current permission for that operation’s target before returning a result; a person approves or dismisses apropose_*proposal in the dashboard, which an external client reads withproposal_detail.operator:act— set a version-fenced triage state.
Use propose_directive with name, condition, action, priority, and excludes when an operator requires exact directive text. A new directive can omit intent only when name, condition, and action are present; otherwise the coach completes the missing fields. Supplied fields remain verbatim, and an edit keeps every field the request omits. priority ranges from 0 through 100. Each excludes entry must match a built-in or authored directive on that agent.
For a reviewed operation, the external client shows the exact result before it calls the write-scoped execution tool. Draft, reversible changes use conversational confirmation in the client. Changes that go live, cannot be undone, or use quota require the grant’s own user to approve the exact digest on the signed-in Radioso review page; execution returns approval_required with that page’s link until they do. Radioso binds execution to that review’s digest, grant, client, target fence, and expiry. MCP-reviewed operations stay on that lifecycle; dashboard Ray proposals keep their existing dashboard review and application flow.
Named client setup is marked verified only when its displayed build has a passing exact-build artifact. Other named builds are unavailable in the chooser, and the generic standards-based setup is labeled unverified. The consent screen identifies the actual client and redirect host, warns about accessible workspace data, and lets you choose the workspace, scopes, and offline_access independently.
Start with what needs attention
Ask Ray “what needs my attention?” and it answers in one read instead of opening five sections. The digest is ranked the way an operator works through a morning:
- Someone is waiting. Conversations handed to a person and approval requests a routine paused on, longest wait first. A handoff nobody has picked up shows no owner, so an unassigned wait reads differently from one an operator is already on.
- Something is failing. Documents that failed to process, sources whose last sync did not succeed, eval cases sitting at failing or error, and turns a customer left a written complaint on — most recent first, because for a failure the useful question is what just broke.
- Backlog. Untriaged review counts per quality signal, and how many documents are queued or still processing.
Every line carries a link to the dashboard surface that resolves it, so “open the failed one” lands on that document rather than on the Knowledge Base index. A conversation that is already an approval or a handoff does not also appear as a review line; the escalation is the more urgent statement of the same work.
Ray also reports which of its six reads it managed. A section your session may not read, or one that errored, comes back marked rather than empty — so “nothing in Quality” and “Quality could not be read” never look the same. Members hold no quality permission, for example, and their digest says so instead of reporting a clear backlog.
Name an agent to narrow it: “what needs attention on the Support agent?” scopes the conversations, approvals, quality signals, and eval cases to that agent. The knowledge base stays workspace-wide, because every agent answers from it.
Each source contributes at most ten lines, and reports the number of rows it matched alongside them. A workspace with 60 failing eval cases sees ten of them and the count 60. Documents and document sources are counted separately, so a run of failed documents cannot push the broken sync that caused them out of the digest. The backlog counts stand for whole groups rather than for individual rows, so they are always complete.
Work the queue
The digest orients a session; the queue is what you work through. Ask “show me everything waiting on a person” and Ray lists the three kinds where the next move is yours — approval requests a routine paused on, conversations handed to a person, and turns a customer left a written complaint on — longest wait first across all three, and no source is capped in favour of another. Narrow it with “just the handoffs” or “only the Support agent”, and raise the page with “show me fifty”. Complaints are read from the end of their own queue, so the oldest are exact at any page size; handoffs are ranked over a window ten times the page you asked for, so one nobody has touched in months can sit outside it. Each source reports how many rows it matched, so a short page reads differently from an empty queue.
Each row carries what its follow-up needs: the decision behind an approval, the turn behind a complaint and the version that turn was read at, and who holds a handoff (or nobody, for one still waiting). That is what lets the next sentence be an instruction rather than a second search.
Once you have decided what a turn was, tell Ray to record it: “mark that one resolved — it was a knowledge gap, I added the shipping page.” Ray writes the triage state and the resolution reason against the version it read. If another operator moved the same row while you were talking, Ray tells you what they set and leaves their decision alone, so two people working the queue at once never overwrite each other. Nothing a customer sees changes either way.
For the reply itself, ask “draft a reply for that conversation.” Ray replays the agent over that conversation’s own transcript, with the agent’s current instructions, directives, skills, and knowledge, and hands you the text it would send. The run is ephemeral — the draft exists in your Ray thread and nowhere else — so you read it, edit what you want, and send it yourself from the conversation. It answers the customer’s most recent message even when the agent replied after it — which is what a handoff and a complaint both look like — and needs the conversation to have a customer message at all.
If the conversation is part-way through a routine, the draft resumes from where it stands rather than starting over, and says whether it did. A conversation paused on a pending approval is the exception: the agent itself would not take a routine step while it waits on your decision, so neither does the draft. And a message you have already answered by hand has nothing outstanding to draft for, so Ray says so instead of handing you a duplicate.
The drafting turn is held to the same rule as a test turn: it answers, and it does nothing else. Every skill that reaches outside the conversation — a notification, a webhook, a contact send, an external tool — stays suppressed, so composing a draft cannot act on the customer’s behalf. Context variables resolve through the same outside path, so a draft for an agent whose instructions interpolate them renders without their values.
Where Ray stops. Ray drafts; you send. Taking a conversation over, handing it back, transferring it to a colleague, and approving or refusing a pending decision are yours as well — ask for any of them and Ray says so and gives you the link. Deciding who is answerable to a waiting customer is a person’s call, and that is the point of the line.
Ask about what you see
The panel captures the dashboard context that is already on screen. A selected conversation row, agent, routine, eval case, or trace stage can be included in the next turn, with display names shown in the context rail.
To ask about a particular sentence, select the text in the dashboard. Choose Ask Ray beside the selection. Ray opens with the text quoted in the composer and sends it as selected context, limited to 2,000 characters.
When a turn reads a specific conversation, agent, or routine, the answer shows a Read during this turn chip. Choose a chip to open the matching dashboard surface. If the current screen has no display name for an entity, the chip uses its id so the link remains useful.
On an empty thread, suggested questions adapt to the view and focused entity. For example, a focused conversation offers to explain what happened in that conversation.
Ask how Radioso works
Ray reads Radioso’s own documentation, so “what does the priority field on a directive actually do?” is a question you can ask in the panel rather than leave for the docs site. Ray reads the pages published with the release you are running, which is what makes the answer worth trusting on an install that sits a few versions behind the current one.
The useful part is that Ray holds both halves at once. Ask how directive matching works and it can explain the mechanism, then tell you this agent has four directives and two of them never match because a third excludes them.
Ray cites the page it read, so you can open the full version when you want it. Questions about your own uploaded content are answered from the workspace knowledge base, so “how do directives work” and “what does our refund policy say” reach different sources.
Review proposals
When Ray identifies a configuration change, it can draft a proposal for a new agent, a directive, an agent setting, a routine, a skill’s configuration, a context variable, a document, a website crawl, or the workspace’s ingestion settings. The proposal card names the target, summarizes the change, and shows the current and proposed values before anything is written. Expand Show changes to inspect the field-level diff.
Applying a proposal for an agent’s custom instructions, directives, routine definitions, or per-agent context selection writes the agent’s private draft. Open Test Chat to exercise that saved candidate, compare it with the published revision, and run selected eval cases before publishing. Sample context values used in a test stay with the test input; publishing never sends them to a visitor. Proposals for shared context-variable definitions, skill configuration, documents, crawls, or workspace settings keep their existing live configuration flow and are reviewed separately from the agent draft.
A routine proposal comes in two shapes: a new routine, and an edit to one that already exists. See Change a routine.
A directive removal proposal targets an existing directive by id and drafts nothing else. The card shows the directive as it stands on the current side and a notice that it will be permanently removed on the proposed side. Applying it deletes the directive, and that cannot be undone.
When a directive is misfiring, Ray can instead draft an enablement proposal. Disabling takes the directive out of play while keeping its condition, action, binding, and other authored settings ready to review or re-enable later. A re-enable proposal validates the directive’s binding again before it can fire.
A skill config proposal creates a new skill or updates an existing one by the same rules the Skills editor enforces: a setting that depends on another stays refused while its parent is off, and a notify skill needs at least a recipient email or a webhook URL before Ray will propose it — Ray asks you for either rather than inventing one. Updating a skill can change its target, configuration, invocation mode, or enabled state, but not its name or capability — Ray proposes a new skill for either of those.
A context variable proposal can create or update the variable’s definition, enable or reconfigure it for the agent, or do both in the same card. A resolver-sourced enablement needs the skill that resolves the value, and that skill must already exist and be enabled on the agent the proposal targets; every other source refuses one. If that skill is disabled or removed before you apply the proposal, the card becomes Stale and writes neither the definition nor the enablement. Applying the proposal writes the definition, the enablement, or both, in that order, so a proposal that creates a variable and enables it for the agent in one step always has an id to enable by the time it runs.
A document proposal covers the three knowledge changes Ray can justify from what it read. It drafts a new document when a turn exposed a gap the workspace has no answer for — Ray writes the whole body, up to 20,000 characters, and the card shows it in full. It drafts a retrieval change when a document exists but reaches answers it should not, or lacks the metadata a rule filters on: retrieval eligibility, an expiry date, and the metadata map, each reversible. And it drafts a removal for a document that should not exist at all, which deletes the document and everything indexed from it and cannot be undone.
Rewriting a document’s text stays with you in the Knowledge editor. Ray reads
documents as search snippets and retrieved chunks, so it sees parts of a document
rather than the whole of one, and a body it drafted as a replacement would lose
whatever it never read. When a document says the wrong thing, ask Ray which
passages are wrong — document_chunks quotes them — and edit from there.
A crawl proposal points at a website. Drafting one costs nothing; applying it fetches the site and indexes what it finds, which spends real crawl budget and reaches an external server, so it goes on a card rather than starting the moment Ray decides it would help. The card names the URL, the page ceiling the deployment allows, and any URL patterns Ray narrowed the crawl to. To refresh a site that is already a source, ask Ray to recrawl it instead — that reuses the crawl settings the source already carries.
An ingestion settings proposal changes how documents are chunked and enriched when they are processed: the chunking strategy, the fixed-window size and overlap, the structured minimum and maximum, and the enrichment switches. Name the one field you want changed and the rest carry over from the stored settings, because the write replaces all of them at once — the card shows the whole result so nothing changes out of sight. Applying it re-chunks nothing on its own; reprocess a document or a source afterwards to reach what is already indexed.
The embedding model is a cost decision with a typed confirmation of its own in Knowledge → Ingestion: changing it re-embeds every chunk in the workspace. Ray describes the change and hands you the link.
A workspace settings proposal covers the assistant’s own wording and the public channels it answers on: the assistant name, greeting, default locale, custom instruction, and suggested questions, plus the anonymous chat link, the website embed, its allowed origins, and its launcher label and position. Name the one field you want changed and the rest carry over from the stored settings, the same way ingestion settings work, because this write also replaces all of them at once.
Some of those fields decide who can reach the agent rather than what it says — turning on the anonymous chat link, turning on the embed, or adding an allowed origin. A proposal that changes one of them says Changes who can reach the agent on the card, and the Apply confirmation asks that question instead of the usual one. It is a reversible change, and the card tells you which kind of decision you are making before you expand the diff.
Applying a workspace settings proposal goes through the same service the Settings → Workspace page writes through, so it gets the same validation and records the same audit events an operator’s own edit does. The anonymous chat and website embed tokens stay with the operator: Ray reads neither, and applying a proposal rotates neither. Ask Ray to change the behavior of one agent among several and it drafts an agent setting proposal instead, which targets that agent by name.
Apply asks for the permission that governs what the proposal changes, not one
permission for all of them: workspace.agents.manage for an agent, directive, agent
setting, routine, skill, or context variable; workspace.documents.manage for a
document or a crawl; workspace.settings.manage for ingestion settings and
workspace settings. An
operator who manages knowledge but not agents sees Apply on the document cards
and a read-only view of the rest. Ray sends the confirmed proposal through the same guarded
management service used by the dashboard. Operators with read access can
review or dismiss a pending proposal; the Apply action is available only when
their session can manage agents.
A proposal that Ray measured first says so. The card states what it was verified against — “Verified against 3 cases — 2 improved, 1 regressed” — and expanding it lists each case with the verdict it recorded before and the verdict the proposed change produced. Regressions are counted alongside fixes, because the decision the card exists to support is whether the trade is worth making. A proposal Ray drafted without replaying anything carries no evidence section, so an unmeasured change never reads as a verified one.
Each proposal is checked against the configuration it changes. A change to an unrelated setting does not block it. If one of its own fields changed, the card becomes Stale, names that field, and leaves the configuration untouched. If the target was removed, the card says so. Ask Ray to draft the change again so the proposal reflects the current configuration. Applied proposals link to the target entity, while failed proposals show the reason returned by the management service.
Create an agent
Ask Ray to build an agent for a site and it reads the site first. Give it a URL — “set up an agent for https://acme.example.com ” — and it fetches the landing page and a sample of what that page links to, then comes back with a name, an answering instruction, a greeting, a locale, the contact and privacy links it found, and a chunking strategy with the reasoning behind the choice. Nothing is created at this point, and nothing is stored.
Reading a site costs real time and reaches an external server, so the analysis is rate-limited against the same budget as Ray’s other expensive work. A site that refuses automated visitors, needs a login, or carries too little text comes back as a plain failure rather than a guess.
You can take the suggestion, change any part of it in conversation — a different name, a stricter instruction, a locale the site does not declare — and Ray drafts the agent as a proposal. The card names the agent, the site it will be grounded in, and the reasoning Ray gave for the configuration. Apply creates the agent and queues that site for ingestion, so the agent has something to ground on as soon as the first documents finish processing.
The chunking strategy is the one suggestion the agent card leaves out. Chunking is a workspace-wide ingestion setting, so adopting it re-chunks every source you already have, not just the site you are adding. Ray keeps it in the analysis, and if you want it, asks for it as an ingestion settings proposal of its own.
A proposal Ray never applied leaves nothing behind, which is the point: a first guess at an agent for an unfamiliar site is worth reviewing, and a dismissed card costs less than an agent to clean up.
Applying happens in steps — the agent, then its locale and contact settings, then the website. If a step after the agent fails, the card still reports the agent as created and links to it, and names the step that did not finish. An agent that exists is worth telling you about even when its site did not queue: you can start the crawl from Knowledge, or ask Ray to crawl it, without creating the agent a second time.
Two limits are worth knowing. Ray never deletes an agent — it describes what you would be removing and links you to the agent’s settings, where the confirmation lives. And it always creates from a website: to set one up by hand, choose New agent in the agents list and take the create manually option.
Once an agent exists, changing it is a different proposal — see the agent setting card in Review proposals.
Change a routine
Most routine work is fixing one that already exists. Tell Ray what misbehaves — “in the Order status routine, the first step asks for the order number without saying why” — and it drafts an edit to that step alone.
Ray addresses each change by the element you named: the routine’s name, its trigger, priority, and re-entry mode; the wording of a step or an ending; an information field’s description or whether it’s required. The proposal card shows one row per changed element, so you review the step that changed rather than the whole flow.
Changing the routine’s shape — adding or removing a step, or rewiring which step leads where — happens in the routine editor. Ask Ray for one of those and it points you there.
Applying an edit writes it into the agent’s private draft, next to the directives, skills, and context variables you have changed there. What your agent is serving keeps running until you Review & Publish the agent.
Ray checks an edit before it becomes a card. If the change would introduce a validation problem the routine doesn’t already have — a step referring to an information field that isn’t defined, say — Ray tells you what would break rather than drafting it. You can ask for that check on its own too: “does the Order status routine validate?” reports each diagnostic and where it is.
Once the edit is applied, a good order is: test it with a real message, replay the eval cases it touches, then Review & Publish the agent.
Test a routine change
Ray can run a representative message through the selected agent before you publish the change. For example, ask: “Test the Support agent with ‘I need to return a damaged item’ against the Returns routine. Did it trigger?”
The test uses the same routing, retrieval, directive, and routine pipeline as a normal agent turn. Ray receives the agent’s answer, citation references, the turn outcome, stable conversation and message ids, and a sanitized trace of the stages that ran. The result links to its test conversation. Its serialized payload is capped at 32,000 bytes; when the cap removes trace stages, citations, or answer text, the result identifies those omissions.
This turn runs with a probe effect policy. Retrieval can supply evidence, and the probe conversation keeps its own routine, pending-decision, clarification, and directive state so a follow-up can continue the draft flow. External skill delivery, queued routine actions, ownership handoffs, customer analytics, and conversation summaries stay untouched. The test request, answer, and audit record are saved as operator-test traffic, outside the normal customer Inbox, Quality, and Audience Pulse populations. A follow-up can continue only the probe conversation created for the same Ray thread, operator, workspace, and agent.
Testing requires all four permissions: workspace.agents.read,
workspace.chat.use, workspace.history.read, and
workspace.agents.manage. The rest of Ray remains available with
workspace.agents.read; the test-turn capability appears only when the session
holds the complete set.
Agent-turn testing is a dashboard Ray capability. The standalone MCP server does not expose it.
Capture and re-run eval cases
When Ray finds a turn that answered badly, ask it to keep the turn: “capture that turn as an eval case.” Ray freezes the conversation and the settings that produced the answer, then links you to the case so you can add the expectations that should hold. Capturing the same turn twice opens the case that already exists, and Ray says which of the two happened.
Once you have cases, Ray can re-run up to five of them in one request and report each outcome next to the whole suite’s pass rate: “re-run the refund cases and tell me what moved.” Ask Ray to list the cases first, then name the ones your change should have affected. The limit is deliberate — cases replay one after another and each full replay costs an answer call and a grading call, so a whole library behind a single request would leave you waiting on it. Ray runs the full agent pipeline; ask for a retrieval-only run when you only want to see what retrieval found.
A case with no expectations comes back as skipped, because there is nothing to score until you add an assertion. Failing cases come back with the assertions that failed and the reason each one gave. If a case id does not exist in the workspace, Ray names it rather than leaving you to read its absence as a pass.
Try a change before you propose it
Ray replays a single case against a configuration that is not live yet, so a proposal arrives with a measured result: “replay the refund case with this instruction and tell me whether it passes.” The replay accepts a custom instruction, a greeting instruction, directives, skill settings, a different model, retrieval settings, and a mid-routine starting position — every setting that changes how the agent behaves — and reports the verdict that configuration produced, the answer, the grounding verdict, the assertions that still fail, and the model that answered. A replay that never produced an answer comes back as an error carrying the reason, so a failed turn reads as a failure rather than as a case with nothing to score.
A replay leaves the library where it was. The case keeps the verdict it recorded, its last run stays the one you see in the Eval section, and the suite’s pass rate holds, so trying five variations of an instruction costs five answers and moves nothing. Ray reports the case’s recorded verdict next to the replay’s own verdict, which is the comparison a proposal rests on.
Each replay runs a full agent turn and costs one, so name the case a change should move rather than the library. When you decide to keep a result, re-run the case with the suite so the library records it.
Ask Ray to propose the change it just measured and the proposal carries those replays with it. The measurements come from the runs Ray recorded, not from its account of them, and each one is tied to the agent, the conversation, and the change it was measured on: a replay of one change cannot be attached to a proposal for a different one. A setting proposal has to cite a replay that put that exact value under test, and a directive proposal has to cite one that ran with directives in place.
A directive removal or disable proposal cites replay evidence the opposite way: ask Ray to replay a case with that specific directive excluded. The replay service resolves the directive against the agent’s real configuration and drops it before running the turn, so the exclusion it records is something Ray cannot fake by handing back a directive list that merely looks like it left the directive out. Cite a replay that did not explicitly exclude the directive — including one built by hand-editing the directive list — and Ray refuses it, because nothing proves that replay ran without the directive in force.
A replay can exclude several directives at once — a fine way to try “what if I dropped both of these” — but that evidence does not back a proposal to affect one directive by itself. A proposal to affect one directive by itself needs a replay that excluded only that directive: a replay that dropped it alongside another one measured a configuration where both were gone, and removing two directives together can move a case’s verdict in a way that removing either one alone would not. Ray refuses to cite that replay for a single-directive removal or disable and asks for one that excludes just the directive being proposed.
That replay also has to be about nothing else. One that also tried a different model, different instructions, different retrieval settings, or a routine starting position measured that whole combination, not the directive’s absence on its own, so a passing verdict could be down to any of those other changes instead. Ray refuses to cite a replay like that for a directive removal or disable and asks for one where excluding the directive is the only override in play. An exclusion replay measures the directive absent from the run, so it cannot support a proposal to re-enable it.
Routine proposals do not carry replay evidence. No replay override installs a routine, so a passing replay would say nothing about the routine being proposed — test the routine with a real turn instead.
A skill config proposal can carry replay evidence only for an agent’s default retrieve skill — the one that answers without a routine naming it. A replay can put that skill’s settings under test the same way it can a setting proposal’s value; every other skill capability has no equivalent replay override, so Ray proposes those changes unmeasured.
A context variable proposal never carries replay evidence. A replay measures the configuration Ray sends the model — instructions, directives, skill settings — and none of the settings a replay can override install a pushed, browser, or resolver value for a visitor session, so no replay can speak to a context variable proposal either way.
A replay runs against the configuration the eval case froze, never the agent’s current one, so a measurement is dated by any edit to that agent after the case was captured — whether the edit landed before the replay or after it. When that has happened the card marks those measurements as describing an earlier configuration rather than dropping them, so you can see both what was measured and that it has aged. Capture a fresh case from a recent turn when you want a measurement against today’s agent.
These capabilities need workspace.retrieval.query, the same permission the
Eval section uses.
Follow a running turn
While Ray works, activity lines show the UI-safe capability label and its stage. For example, Reading conversation trace — started becomes completed when that read finishes. After the turn completes, the timeline collapses to a one-line summary; choose it to inspect the individual reads. A failed activity stays visible so you can distinguish an incomplete investigation from a successful one.
The answer carries a terminal outcome:
- Completed means the turn reached its normal end.
- Budget exhausted means Ray returned its partial findings after its runtime budget ran out.
- Turn failed means the turn could not complete. Read the visible activity and error state before retrying. Choose Retry to send the same question again with the current page context.
Only one turn runs in a Ray conversation at a time. The composer is disabled while a turn is active, and a conversation that is already running shows the same locked state when you open it.
Manage conversations
Your Ray conversations are separate from customer Inbox, quality, and history views. Return to Ray to resume a saved thread. Use Delete in the thread header and confirm to remove the selected conversation and its messages.
Ray is available to operators with workspace.agents.read. It uses the
dashboard session, so personal and service API credentials cannot start a Ray turn.
When the workspace has no resolved LLM capability, Ray shows a provider
setup state with a link to Settings → Providers.
A Ray conversation is kept for 90 days after its last message, then removed
along with its messages and proposals. Threads you want to keep past that are
worth summarizing somewhere durable — a spec, an issue, or the rationale field
on the change you applied. Your operator can change the window with
COPILOT_CONVERSATION_RETENTION_DAYS.
Every change Ray makes is recorded with the operator who applied it and the surface they applied it from, so a configuration change that surprises someone can be traced back to the person and the session that made it.
Budget for verification
Testing a turn, replaying a case, drafting a reply, and running a suite each cost a real model call — that is what makes their evidence worth anything. Ray spends at most six replayed turns in one Ray turn, and a suite run counts once per case, so asking for five cases spends five. Past the budget it stops, answers with what it measured, and asks you to send another turn.
In practice this shapes how to ask. “Replay the three refund cases and tell me which regressed” fits in one turn with room to spare. “Replay everything in the library” does not, and Ray will get part of the way and stop. Ask for the cases the change should have moved, read the result, then ask for the next batch.
Common failure modes
- The Ray entry is missing — your session does not hold
workspace.agents.readfor the active workspace. - Provider setup appears — configure or select an LLM under Settings → Providers, then return to Ray.
- A turn remains locked — another turn is still running in that conversation. Select a different Ray conversation or wait for the running turn to finish.
- A tool activity fails — Ray continues where it can and marks the affected turn honestly. The answer identifies the data it could not access.
- Document maintenance is unavailable — your session can read documents
but does not hold
workspace.documents.manage. Chunk inspection remains available; reprocess and recrawl do not appear in Ray’s tool set. - Retrieval probing is unavailable — your session holds
workspace.retrieval.querybut notworkspace.agents.read, which the probe needs in order to run as an agent. - A source cannot be recrawled — recrawl accepts an existing website source with a stored URL. Uploaded, API, and connector sources use their own refresh paths in the Knowledge Base.
- A proposal is stale — the card names the changed field or says that the target was removed. Ask Ray to prepare the change again and review the updated diff.
- Apply or Dismiss says conflict — another apply on that proposal is already running, or one never finished. Wait a few minutes and try again; an apply that never finished releases its hold on its own, after which the proposal can be applied or dismissed.
- An apply reports that an earlier attempt may already have taken effect — an apply of that card stopped partway, and the card creates something new: a document, or a crawl. A second attempt would create a second one, so Ray resolves the card instead of retrying it. Open the Knowledge Base to see whether the document or crawl is already there, then ask Ray for the change again if it is not. Cards that change something already in the workspace — a directive, a setting, a document’s retrieval state — retry safely and come back Stale if the target moved.
- Agent-turn testing is unavailable — your session is missing one of
workspace.chat.use,workspace.history.read, orworkspace.agents.managein addition to Ray’s agent-read permission. - A routine cannot be tested — confirm that the routine belongs to the selected agent.
- A suite run is refused as too large — one request runs at most five cases. Ask for the cases the change should have moved, then run the rest in a second request.
- Ray stops verifying mid-answer — the turn spent its budget of six replayed turns. Read what it measured, then send a second turn for the rest.
- Verification is rate limited — too many test turns, replays, or suite runs in a short window. Ray says how long to wait and answers from what it already has.
- A Ray conversation is gone — it passed the 90-day retention window. Deleted conversations are not recoverable.
- A test conversation cannot be continued — start a fresh test from the current Ray thread. Probe conversations are scoped to the Ray thread, operator, workspace, and agent that created them.
Read next
- Human takeover — investigate a customer conversation and take ownership when a person needs to reply.
- Test your agent in the workbench — run an agent turn directly and inspect its customer-facing trace.
- Evaluate answer quality — turn a known behavior into a repeatable quality check, and read what a case’s assertions can check.