> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Screen share

> Let a user share their screen and talk to the agent about what is on it.

In a screen session, the browser sends two things on one WebSocket: images of the shared screen, and microphone audio. The agent answers out loud about what it sees.

<Tip>
  If you only need an agent on your site, the [embed widget](/embed-widget) already supports screen share.
</Tip>

## Before you start

* An agent. Screen share is on by default.
* An API key.
* A [token endpoint](/realtime/browser-voice) on your server.
* An HTTPS page, or `localhost`. Screen capture does not work on plain HTTP.

<Warning>
  Three mistakes break most screen-share apps. None of them shows an error. The socket stays open and the agent never speaks.

  1. **Screen capture does not include the microphone.** The audio from `getDisplayMedia` is the computer's sound, not the user's voice. You must ask for the microphone separately.
  2. **Images alone never get a reply.** The agent only takes a turn when it hears speech.
  3. **The browser may ignore your audio sample rate.** The audio then arrives at the wrong speed and the agent cannot hear it.

  The client library handles all three.
</Warning>

## Build it with the library

<Steps>
  <Step title="Create screen sessions on your server">
    Use the token endpoint from [Voice in the browser](/realtime/browser-voice). The library adds `?source=screen` to the request, and the endpoint passes it on.
  </Step>

  <Step title="Start the session">
    Pass `source: "screen"` and a `<video>` element for the preview. The library opens the screen picker, asks for the microphone, and merges the two.

    ```js theme={"system"}
    await session.start({ source: "screen", videoElement: preview });
    ```
  </Step>
</Steps>

Here is the whole app:

```html screen.html theme={"system"}
<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8" />
    <title>Screen share</title>
  </head>
  <body>
    <button id="start">Share screen</button>
    <button id="stop" disabled>Stop</button>
    <p id="status">idle</p>
    <video id="preview" muted playsinline style="width: 100%"></video>
    <div id="transcript"></div>

    <script type="module">
      import { EclatiraSession } from "https://cdn.jsdelivr.net/npm/@eclatira/realtime@0.1.3/eclatira-realtime.js";

      const startButton = document.getElementById("start");
      const stopButton = document.getElementById("stop");
      const status = document.getElementById("status");
      const preview = document.getElementById("preview");
      const transcript = document.getElementById("transcript");

      let session = null;
      let line = null;
      let speaker = null;

      function render(who, text, final) {
        // Transcripts arrive as revisions of the current utterance, not as
        // deltas, so replace the live line instead of appending fragments.
        if (who !== speaker || !line) {
          line = document.createElement("p");
          transcript.appendChild(line);
          speaker = who;
        }
        line.textContent = `${who}: ${text}`;
        if (final) speaker = null;
      }

      startButton.onclick = async () => {
        startButton.disabled = true;
        session = new EclatiraSession({ tokenUrl: "/api/eclatira/token" });

        session.on("transcript", ({ speaker, text, final }) => render(speaker, text, final));
        session.on("speakingchange", ({ user, agent }) => {
          status.textContent = agent ? "agent speaking" : user ? "listening" : "live";
        });
        session.on("error", (error) => render("error", `${error.code}: ${error.message}`, true));
        session.on("close", ({ code, reason }) => {
          status.textContent = `closed (${code}${reason ? ` ${reason}` : ""})`;
        });
        session.on("stopped", () => {
          startButton.disabled = false;
          stopButton.disabled = true;
        });

        try {
          await session.start({ source: "screen", videoElement: preview });
          status.textContent = "live";
          stopButton.disabled = false;
        } catch (error) {
          status.textContent = `failed: ${error.message}`;
          startButton.disabled = false;
        }
      };

      stopButton.onclick = () => session?.stop();
    </script>
  </body>
</html>
```

The library also watches for the browser's own **Stop sharing** button. When the user clicks it, the session ends.

To ask the agent about the screen without speaking, call `session.sendText("What am I looking at?")`. The agent replies right away.

## Build it without the library

Read [Media format](/realtime/media-format) first. The audio rules are strict, and mistakes fail with no error.

<Steps>
  <Step title="Get the screen and the microphone">
    Ask for the screen with `audio: false`, then ask for the microphone separately. Merge them into one stream.

    ```js theme={"system"}
    // Screen video. audio:false is deliberate — display audio is not the mic.
    const display = await navigator.mediaDevices.getDisplayMedia({
      video: { width: { ideal: 1280 }, height: { ideal: 720 } },
      audio: false,
    });

    // The microphone, separately.
    const mic = await navigator.mediaDevices.getUserMedia({ audio: true, video: false });

    // One stream carrying both, which is what the rest of the pipeline consumes.
    const stream = new MediaStream([...display.getVideoTracks(), ...mic.getAudioTracks()]);

    // The browser's own "Stop sharing" control ends the track without touching
    // the socket, so end the session deliberately when it fires.
    display.getVideoTracks()[0].addEventListener("ended", () => teardown());
    ```

    If the user blocks the microphone, stop and tell them. A session without it never replies.
  </Step>

  <Step title="Create the session, then connect">
    Get the media **first**, then create the session. The screen picker can stay open for a long time, and the token expires after 60 seconds. If you create the session first, it may expire before the user picks a window.

    ```js theme={"system"}
    const { ws_url } = await fetch("/api/eclatira/token?source=screen").then((r) => r.json());
    const ws = new WebSocket(ws_url);
    await new Promise((resolve, reject) => {
      ws.onopen = resolve;
      ws.onerror = () => reject(new Error("could not open the realtime socket"));
    });
    ```
  </Step>

  <Step title="Send images">
    Draw the video onto a canvas, encode it as JPEG, and send the base64 data. Remove the `data:image/jpeg;base64,` prefix.

    ```js theme={"system"}
    const video = document.getElementById("preview");
    video.srcObject = stream;
    video.muted = true;
    video.playsInline = true;
    await video.play();

    const canvas = document.createElement("canvas");
    const context = canvas.getContext("2d");
    const MAX_EDGE = 1280;

    function sendFrame() {
      if (!video.videoWidth || video.readyState < 2) return;

      // Screen shares are often 2560px or wider. The agent reads the frame as an
      // image rather than scrolling it, so downscaling costs nothing legible and
      // saves a lot of bandwidth.
      const scale = Math.min(1, MAX_EDGE / Math.max(video.videoWidth, video.videoHeight));
      canvas.width = Math.round(video.videoWidth * scale);
      canvas.height = Math.round(video.videoHeight * scale);
      context.drawImage(video, 0, 0, canvas.width, canvas.height);

      const dataUrl = canvas.toDataURL("image/jpeg", 0.8);
      ws.send(
        JSON.stringify({
          mime_type: "image/jpeg",
          data: dataUrl.slice(dataUrl.indexOf(",") + 1),
        })
      );
    }

    sendFrame();
    const frameTimer = setInterval(sendFrame, 2000);
    ```
  </Step>

  <Step title="Send audio">
    Send microphone audio on the same socket. The order of audio and images does not matter. The audio code is the same as a voice session. See [Voice in the browser](/realtime/browser-voice#build-it-without-the-library).
  </Step>

  <Step title="Handle messages">
    ```js theme={"system"}
    ws.onmessage = (event) => {
      const message = JSON.parse(event.data);

      // Typed messages carry `type`; voice turn and content events do not. Check
      // `type` first, then fall back to field presence.
      if (message.type === "error") {
        console.error(message.code, message.error);
        return;
      }
      if (message.type === "session_created") {
        console.log("session", message.session_id);
        return;
      }

      if (message.interrupted) flushPlayback();

      for (const part of message.parts ?? []) {
        if (part.type === "audio/pcm") enqueue(part.data); // base64, 24 kHz PCM
      }
      if (message.output_transcription) render("agent", message.output_transcription.text);
      if (message.input_transcription) render("user", message.input_transcription.text);
    };
    ```
  </Step>

  <Step title="Clean up">
    Stop **every** stream: the merged one, the screen and the microphone. Otherwise the browser's sharing bar stays up.

    ```js theme={"system"}
    function teardown() {
      clearInterval(frameTimer);
      ws.close(1000);
      [stream, display, mic].forEach((s) => s?.getTracks().forEach((t) => t.stop()));
      // Also close the two AudioContexts from the microphone pipeline.
      recorderContext?.close();
      playerContext?.close();
    }
    ```
  </Step>
</Steps>

## How often the agent sees the screen

The server passes **at most one image per second** to the agent. If you send more, it keeps the newest one and drops the rest.

* **Sending extra images is safe.** It costs bandwidth, but the agent always gets the latest screen.
* **Send one about every second** for the freshest view. The library sends one every 2 seconds by default. Set `frameIntervalMs: 1000` to match the server.

The agent sees a series of still images, not video. It can read code, a dashboard or a form. It cannot follow a moving cursor or an animation. After the user switches tabs or scrolls, wait a few seconds before asking about the new screen.

<Note>
  The server has no limit on image size or quality. The only hard limit is 16 MiB per WebSocket message. A longest side of 1280 px at quality 0.8 keeps text readable and images small.
</Note>

## Errors you are likely to see

| Code | Meaning |
| - | - |
| `no_audio_received` | Images have arrived for about 8 seconds with no audio. The agent cannot reply. You probably used the screen's audio instead of the microphone, or the user blocked the microphone. |
| `invalid_base64` | The image data is not valid base64. Usually the data URL was cut off before its comma. |

The session stays open after both. See [Errors and close codes](/realtime/errors) for the full list.

## Turn screen share on or off

Screen share is on by default. Change it with `enable_screen_share`:

```bash theme={"system"}
curl -X PATCH https://app.eclatira.com/api/v1/agents/AGENT_ID \
  -H "Authorization: Bearer ek_your_api_key_here" \
  -H "Content-Type: application/json" \
  -d '{"realtime": {"enable_screen_share": false}}'
```

This setting also controls the embed widget's **Share Screen** button.

When it is off, creating a `screen` session fails with `403` and code `screen_share_disabled`. A token created before the change closes with code `4003` when it connects.

## Keep a record

Walkthrough and training apps often need the conversation afterwards. Save each final transcript to your own storage as it arrives. Do not wait until the end, or a closed tab loses the session. You can also register a [webhook](/webhooks) to get the full transcript when the session ends.

## Next steps

<CardGroup cols={2}>
  <Card title="Media format" icon="waveform-lines" href="/realtime/media-format">
    The exact audio and image formats.
  </Card>

  <Card title="Troubleshooting" icon="bug" href="/realtime/troubleshooting">
    The agent connects but never speaks.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.