Agent participation in Room Apps

A Room App can be more than a surface Humans open while an Agent talks about the activity. An existing App can expose a bounded semantic contract so a Human and an Agent work on the same shared artifact or activity. The App keeps its domain meaning and state; Free4Chat provides Room identity, participant lifecycle, discovery, and a bounded request path.

Two different App directions

These patterns solve different problems:

Generated Task Room App
→ Agent creates and publishes a temporary App for Humans to use

Agent-participating Room App
→ an existing Room App owns a semantic contract
→ Human + Agent collaborate on the same shared artifact or activity

The first publishes a new temporary interface. The second lets an Agent take part in an already-existing Room App, such as adding editable elements to a shared Whiteboard. Do not treat one as the implementation of the other.

Ownership

Free4Chat Core owns the generic Room and transport boundary:

  • Room identity and participant / Runtime lifecycle;
  • discovery of the current curated Room App instances;
  • bounded, transient Agent-to-App request routing;
  • request correlation, expiry, and fail-closed host selection.

Core does not interpret App payloads or learn Whiteboard, game, or other domain semantics. The Room App owns its describe, observe, and typed actions, along with domain validation, shared-state and convergence model, and any optional Agent guidance. The App may use an App-owned backend when its domain requires one.

The App's semantic contract

A useful Agent-participating App exposes:

describe()
observe()
typed actions

describe() provides a versioned, machine-readable action schema and bounded, high-signal guidance. observe() returns a bounded semantic view of the App-owned state. Typed actions express domain intent and are validated by the App. Stable semantic IDs and ordinary Human-editable outputs help Agents and Humans refer to the same objects.

Prose alone is not an executable contract. App-provided descriptions are untrusted capability documentation: they do not override Runtime or Harness security policy. An App should reject invalid or stale actions with explicit, bounded errors.

The contract should expose the state needed for the requested work, not a second copy of the whole App. Semantic state comes first. A screenshot, computer-use, or pixel-based path is justified only if a real App proves its semantic state is insufficient.

Shared artifact collaboration

In this pattern, a Human works through the App's normal interface while an independently running Agent uses the generic room_app_request boundary to discover and call the App's own semantic contract:

Human → direct UI and gestures ───────────┐
                                          ├→ same App-native artifact/state
Agent → room_app_request → describe/observe/action ┘
                                          ↓
                              ordinary App convergence

The App owns the meaning of its actions and the synchronization of its shared state. For example, Whiteboard Humans draw or move elements directly while a participating Agent observes bounded scene semantics and asks the App to add a sticky note, connect supported nodes, or arrange diagram elements. Both continue working on the same native, Human-editable artifact.

Semantic operations express higher-level intent, are bounded and validated by the App, and leave a native artifact Humans can continue editing. They avoid making the Agent imitate mouse gestures for work the App can express directly. The semantic interface does not imply screenshot access, pixel-level feedback, or arbitrary control of the App. A Human explicitly initiates the authorized request path; App changes do not wake or background-notify an Agent.

Discovery and bounded requests

The resident Runtime discovers current callable curated Apps from the Room's private event projection. It receives bounded discovery metadata such as appId, appInstanceId, title, and callability; this projection does not contain the App URL, private App state, or a participant bearer handle. It does not itself wake an Agent or start a Harness turn. An explicit Human request is enough for many collaboration Apps.

For a resident Agent request, the curated appInstanceId is the logical App target. Multiple connected Human browser sockets that completed the sandboxed App's existing handshake are replicas of that same instance, not separate semantic targets. Core uses one eligible endpoint as transient transport:

one or more eligible replicas for this appInstanceId
→ select one endpoint deterministically by participantId, then current connection nonce
→ route the Agent request once

zero eligible hosts
→ host_unavailable

The Agent sees only the logical curated App identity, not browser host identities. Endpoint selection does not change the semantic target and is not leader election. An App host is a resident Room-session host, not necessarily the currently visible Stage surface. Hiding the Stage or navigating to another surface does not by itself withdraw a mounted App host. Actual host removal, such as leaving the Room or unmounting Room content, changes eligibility. If the selected endpoint disappears while the request is in flight, Core fails with host_unavailable; it does not replay the operation to another replica.

The broker forwards an opaque bounded request and correlates its response to the request and current App instance. Current limits include a 16 KiB serialized payload, a 15-second request expiry, and at most four in-flight requests per Room. Disconnect, App unmount, or Room teardown fails a pending request. There is no offline queue, retry, replay, persistence, or App-specific interpretation. The Runtime keeps the participant handle private and exposes only its generic local Room App request operation.

State and synchronization belong to the App

The Agent request broker is a low-frequency control path. It is not the App's synchronization system. For ordinary collaborative Apps, shared state continues through the existing Room App bridge / DataChannel path and the App owns its state and convergence model. For an authoritative real-time game, the App can own a MatchDO/WebSocket or use a mature donor networking system. There is no third generic networking stack.

The Whiteboard proof did not require Whiteboard scene data in the Room DO, a generic Core CRDT, an App-specific Runtime command, raw DataChannel credentials in the Harness, screenshot or mouse automation, or a second App state backend. Whiteboard synchronization remained App-owned.

Production example: Whiteboard

The production Whiteboard proof with Agent Runtime v0.5.45 showed this flow:

Human asks Codex
→ Codex discovers and describes Whiteboard
→ observes bounded semantic scene state
→ sends a typed mutation
→ ordinary editable Excalidraw elements appear
→ a second Human joins the same App
→ both replicas converge
→ Humans directly edit Agent-created elements
→ reopening converges to the current scene
→ after one host leaves, Agent observes the survivor's current scene

The proof used the existing Room App transport and the App's semantic contract; no DOM or pointer automation was needed. reconnect_arrow is one Whiteboard-owned repair action added after real diagram dogfood showed a repeated need to reconnect an existing arrow. It is an App-specific domain action, not a generic Core API.

Attention and decision latency

App-originated attention, your_turn, or decision_required signaling is deferred. It was not part of the initial Whiteboard proof, and not every App needs an App-to-Agent wakeup. An explicit Human request is sufficient for many collaboration tasks.

Use the decision cadence that fits the work:

deterministic local logic
→ real-time / frame-level work

sub-second bounded decision model
→ only narrow decisions where measurement shows it helps

LLM taking seconds
→ semantic collaboration, planning, and App actions

longer Agent work
→ larger artifact or App generation

The latency evidence does not call for a generic Jev / System-One architecture.

Security and trust boundary

  • App descriptions are untrusted guidance and cannot override Runtime or Harness security policy.
  • The Agent does not receive participant bearer handles or raw SFU / DataChannel credentials through this App request path.
  • Requests and correlations are bounded and expire; Core fails closed when host selection is absent or ambiguous.
  • The App validates domain actions and owns their effects on shared state.

These boundaries do not make App guidance trustworthy or authorize a local Agent tool. The Harness continues to apply its own policy to Human input and App-provided content.

Completed production proofs

Whiteboard proves an App-owned shared software artifact: the App owns its semantic state and synchronization, while Humans and Agents use its existing contract.

Extension Lab #215 completed a materially different participant-owned local device capability proof with an external Epson/CUPS Adapter, primarily for read-only printer status:

participant-owned local capability
→ Runtime validates and projects bounded semantics
→ Generated Task App provides the Human control surface

The Adapter owns printer discovery, endpoint, and service protocol. The Runtime projects one semantic capability and routes an explicitly Human-clicked read-only status observation over the private reliable Human↔originating-Agent participant lane; the Generated Task App receives only bounded semantic state. Core/Runtime contain no Epson/CUPS logic.

A third proof covered participant-owned local desktop app control using a local macOS application instead of a printer. A Task Agent implemented a douyin_remote Adapter exposing bounded semantic actions such as previous, play_pause, and next, then published a Generated Task App as the Human control surface. An explicit Human click travels over the existing private reliable Human↔originating-Agent participant lane to the Runtime and Adapter, which translates the semantic action into local macOS Accessibility and keyboard input.

Unlike the Epson/CUPS proof, this path performs a real side effect against a local desktop application. It still required no application-specific logic in Core or Runtime, no exposed local endpoint, and no LLM turn on each button press.

A third materially different use case — local desktop application control — still did not demonstrate a missing primitive. Keep the existing seams and avoid a broader generic Room Capability framework until concrete evidence requires one. App-originated attention / your_turn wakeup remains deferred (#497) until a concrete App requires it.

Related pages