> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Camera

> Let the agent see the user's webcam and talk about what it sees.

A camera session is a voice session plus images. The browser sends still JPEG images from the webcam on the same WebSocket as the audio. The agent hears the user and sees what is in front of the camera.

This page covers only the camera parts. The audio side is the same as a voice session. See [Voice in the browser](/realtime/browser-voice) first.

## What it is good for

* **Visual troubleshooting.** The user shows the broken thing and describes the problem.
* **Guided physical tasks.** Assembly, installation or setup, with the user's hands free.
* **Reading documents.** The user holds up a card, a label or a form.
* **Anything that is easier to show than to describe.**

<Note>
  The agent sees up to one still image per second. It is not a video call. The agent cannot follow motion or gestures. See [What the agent can see](#what-the-agent-can-see).
</Note>

## Turn the camera on or off

Camera sessions are **on by default** for every agent. The setting is `enable_webcam`. It is the same setting the dashboard and the embed widget use.

<CodeGroup>
  ```bash Check the setting theme={"system"}
  curl -s https://app.eclatira.com/api/v1/agents/AGENT_ID \
    -H "Authorization: Bearer ek_your_api_key_here" \
    | jq '.realtime'
  ```

  ```bash Turn it off theme={"system"}
  curl -s -X PATCH https://app.eclatira.com/api/v1/agents/AGENT_ID \
    -H "Authorization: Bearer ek_your_api_key_here" \
    -H "Content-Type: application/json" \
    -d '{"realtime": {"enable_webcam": false}}' \
    | jq '.realtime'
  ```
</CodeGroup>

`PATCH` only changes the flags you send. All other settings stay the same.

When the camera is off, creating a `camera` session fails with `403`:

```json theme={"system"}
{
  "error": {
    "type": "permission_error",
    "code": "webcam_disabled",
    "message": "Camera is not enabled for this agent.",
    "param": null,
    "request_id": "9f2c41b7e8d44a0b8c1f6de35a07b214"
  }
}
```

If someone turns the camera off after the session was created but before it connected, the socket closes with code `4003`. Handle both.

## Build a camera page

<Steps>
  <Step title="Create a camera session on your server">
    Use the token endpoint from [Voice in the browser](/realtime/browser-voice). The library adds `?source=camera` to the request, and the endpoint passes it on. If you write your own endpoint, set `source` to `camera`:

    ```js token-endpoint.mjs theme={"system"}
    // Runs on your backend. The API key never leaves this process.
    import { randomUUID } from "node:crypto";

    const minted = await fetch("https://app.eclatira.com/api/v1/realtime/sessions", {
      method: "POST",
      headers: {
        authorization: `Bearer ${process.env.ECLATIRA_API_KEY}`,
        "content-type": "application/json",
      },
      body: JSON.stringify({
        agent_id: process.env.ECLATIRA_AGENT_ID,
        // Unique per concurrent session.
        user_id: `web_${randomUUID()}`,
        source: "camera",
      }),
    });

    const session = await minted.json();
    // Hand session.ws_url to the browser as-is. Do not cache it.
    ```
  </Step>

  <Step title="Start the session with a video element">
    Pass `source: "camera"` and a `<video>` element for the preview:

    ```js theme={"system"}
    await session.start({ source: "camera", videoElement: preview });
    ```
  </Step>
</Steps>

Here is a complete page:

```html camera.html theme={"system"}
<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8" />
    <meta name="viewport" content="width=device-width, initial-scale=1" />
    <title>Eclatira camera session</title>
    <style>
      body {
        margin: 0;
        padding: 24px 16px;
        font: 15px/1.5 ui-sans-serif, system-ui, -apple-system, sans-serif;
      }
      main { max-width: 760px; margin: 0 auto; }
      button { font: inherit; font-weight: 600; padding: 10px 18px; border-radius: 8px; }
      button[disabled] { opacity: 0.45; cursor: not-allowed; }
      #preview {
        width: 100%;
        max-width: 640px;
        background: #000;
        border-radius: 10px;
        margin: 16px 0;
        transform: scaleX(-1);
      }
      #status { margin-left: 12px; font-size: 13px; color: #6b7280; }
      #log { border: 1px solid #e5e7eb; border-radius: 10px; padding: 12px; min-height: 140px; }
      .line { padding: 4px 0; }
      .who { font-weight: 600; margin-right: 6px; }
      .err { color: #dc2626; }
    </style>
  </head>
  <body>
    <main>
      <h1>Show the agent your camera</h1>

      <button id="start">Start camera session</button>
      <button id="stop" disabled>Stop</button>
      <span id="status">idle</span>

      <!-- The preview is mirrored in CSS because users expect a mirror image of
           themselves. The frames sent to the agent are NOT mirrored: the CSS
           transform does not affect what the canvas reads out of the element,
           so text the user holds up stays readable to the model. -->
      <video id="preview" muted playsinline></video>

      <div id="log"></div>
    </main>

    <script type="module">
      // Loaded straight from the CDN mirror of the published npm package, pinned
      // to an exact version. Using a bundler? npm install @eclatira/realtime and
      // import "@eclatira/realtime" instead.
      import { EclatiraSession } from "https://cdn.jsdelivr.net/npm/@eclatira/realtime@0.1.3/eclatira-realtime.js";

      const startButton = document.getElementById("start");
      const stopButton = document.getElementById("stop");
      const statusText = document.getElementById("status");
      const log = document.getElementById("log");
      const preview = document.getElementById("preview");

      let session = null;
      let liveLine = null;
      let liveSpeaker = null;

      function write(speaker, text, isError) {
        // Transcripts arrive as successive revisions of the current utterance,
        // so replace the live line instead of appending a fragment per update.
        if (speaker !== liveSpeaker || !liveLine) {
          liveLine = document.createElement("div");
          liveLine.className = "line";
          liveLine.innerHTML = '<span class="who"></span><span class="text"></span>';
          liveLine.querySelector(".who").textContent = speaker;
          log.appendChild(liveLine);
          liveSpeaker = speaker;
        }
        const span = liveLine.querySelector(".text");
        span.textContent = text;
        if (isError) span.classList.add("err");
      }

      startButton.addEventListener("click", async () => {
        startButton.disabled = true;
        statusText.textContent = "connecting";

        session = new EclatiraSession({
          tokenUrl: "/api/eclatira/token",
        });

        session.on("transcript", ({ speaker, text, final }) => {
          write(speaker, text);
          if (final) liveSpeaker = null;
        });

        session.on("speakingchange", ({ user, agent }) => {
          statusText.textContent = agent ? "agent speaking" : user ? "listening" : "live";
        });

        // Server-side frame rejections and the no-audio warning both arrive
        // here. Surface them; they are the difference between debugging this in
        // a minute and debugging it in a day.
        session.on("error", (error) => write("error", error.code + ": " + error.message, true));

        session.on("close", ({ code, reason }) => {
          statusText.textContent = "closed " + code + (reason ? " " + reason : "");
        });

        session.on("stopped", () => {
          startButton.disabled = false;
          stopButton.disabled = true;
          statusText.textContent = "idle";
        });

        try {
          await session.start({ source: "camera", videoElement: preview });
          statusText.textContent = "live";
          stopButton.disabled = false;
        } catch (error) {
          write("error", error.message, true);
          statusText.textContent = "failed";
          startButton.disabled = false;
        }
      });

      stopButton.addEventListener("click", () => session && session.stop());
    </script>
  </body>
</html>
```

The preview is mirrored with CSS, as users expect. The images sent to the agent are not mirrored, so text the user holds up stays readable.

### Image settings

Set these on the `EclatiraSession` constructor. The defaults work well for most cameras.

| Option | Default | What it controls |
| - | - | - |
| `frameIntervalMs` | `2000` | How often an image is sent, in ms |
| `frameQuality` | `0.8` | JPEG quality |
| `frameMaxEdge` | `1280` | Longest side in pixels. Images are scaled down, never up. |

## The agent needs the microphone too

<Warning>
  **A camera session with no audio never gets a reply.** The agent only takes a turn when it hears speech. Images alone do not start a turn. The session connects, accepts every image, and waits forever.
</Warning>

When images keep arriving with no audio, the server sends this once, after about 8 seconds:

```json theme={"system"}
{
  "type": "error",
  "code": "no_audio_received",
  "error": "Frames are arriving but no audio is. The agent replies to speech, so a camera or screen session must also stream 'audio/pcm' from the microphone (or send a 'text/plain' message) before it will respond."
}
```

The session stays open. Show this error in your UI.

Two things cause it:

1. **No audio track.** You called `getUserMedia({ video: true })` without `audio: true`.
2. **Audio at the wrong sample rate.** The audio arrives, so `no_audio_received` does **not** fire. But the agent cannot understand it. See [Media format](/realtime/media-format#the-audiocontext-sample-rate-problem). The library fixes this for you.

To make the agent comment on an image without the user speaking, send text. The agent replies right away:

```js theme={"system"}
// With the library
session.sendText("What do you see?");

// On a raw socket. text/plain data is plain UTF-8, not base64.
ws.send(JSON.stringify({ mime_type: "text/plain", data: "What do you see?" }));
```

## Handle camera problems

### Permission denied

`getUserMedia` throws an error when it fails. The library asks for the camera **before** it creates the session, so a denied camera does not use up a session. Show a clear message for each case:

```js theme={"system"}
let stream;
try {
  stream = await navigator.mediaDevices.getUserMedia({
    audio: true,
    video: { width: { ideal: 1280 }, height: { ideal: 720 } },
  });
} catch (error) {
  if (error.name === "NotAllowedError") {
    // The user clicked Block, or the page is not on a secure origin, or the
    // browser's own camera permission is off. The user must change it in
    // browser settings; calling getUserMedia again usually fails instantly.
    showMessage("Camera access was blocked. Enable it in your browser settings.");
  } else if (error.name === "NotFoundError") {
    showMessage("No camera was found on this device.");
  } else if (error.name === "NotReadableError") {
    showMessage("The camera is in use by another application.");
  } else {
    showMessage("Could not start the camera: " + error.name);
  }
  return;
}
```

* The camera only works on `https://` pages and `http://localhost`. Testing from a phone against your laptop's IP address over `http://` fails.
* Ask for the camera and the microphone in **one** call. If you ask separately and only the microphone fails, you get a video-only session that never replies.

### The camera stops during the session

A camera can stop on its own. The user unplugs it, another app takes it, or the system revokes access. The socket stays open, but images stop. Listen for it:

```js theme={"system"}
const [videoTrack] = stream.getVideoTracks();
videoTrack.addEventListener("ended", () => {
  // The camera is gone. End the session, or fall back to voice only.
  session.stop();
});
```

### End the session cleanly

`session.stop()` stops the image timer, closes the socket, stops every track and releases the preview. The camera light turns off.

If you build without the library, stop **every track**. Closing the socket is not enough. The camera light stays on until each track is stopped:

```js theme={"system"}
stream.getTracks().forEach((track) => track.stop());
videoElement.srcObject = null;
```

## What the agent can see

The agent gets up to one new still image per second. Each image stands alone. Write your agent's instructions with this in mind.

<CardGroup cols={2}>
  <Card title="Works well" icon="check">
    * Reading text, labels, serial numbers and handwriting held still
    * Identifying objects, parts, damage and colors
    * Checking that a step is done ("hold it up so I can see")
    * Counting a few objects that are not moving
  </Card>

  <Card title="Does not work" icon="xmark">
    * Following motion, gestures or sign language
    * Anything that changes faster than once a second
    * Reading moving or blurry objects
    * Reading small text far from the camera
  </Card>
</CardGroup>

The server tells the agent that the user has shared their camera. It does not tell the agent how often images arrive. Add that to the agent's instructions:

```text Agent instruction theme={"system"}
You receive still images from the user's camera about once a second, not
continuous video. You cannot see motion between frames.

When you need to see something clearly, ask the user to hold it still and steady
in front of the camera, then wait a few seconds before commenting.

If an image is blurry, dark, or cut off, say so and ask for a better view rather
than guessing. Do not describe what you expect to see; describe what is in the
image.

Never claim to have watched something happen. You saw a sequence of stills.
```

## Limits

* **Images are not stored.** You cannot read them back later.
* **The session type is fixed.** You cannot turn on the camera during an `audio` session. Start a new `camera` session instead. The conversation does not carry over.
* **Build without the library?** [Media format](/realtime/media-format#video-frames) has the full image pipeline.

## Next steps

<CardGroup cols={2}>
  <Card title="Screen share" icon="display" href="/realtime/screen-share">
    The same idea, for the user's screen.
  </Card>

  <Card title="Troubleshooting" icon="bug" href="/realtime/troubleshooting">
    The agent connects but never speaks.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.