> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice in the browser

> Build a live voice conversation with an agent in your own web page, using the @eclatira/realtime client library.

This guide builds a working voice conversation in a browser tab. The user can interrupt the agent at any time. Every code sample is a complete file you can copy.

You need three parts:

1. **A token endpoint** on your server. It creates a session with your API key.
2. **The client library**, `@eclatira/realtime`, on your page.
3. **A page** with a start button.

<Info>
  The library handles the hard parts of browser audio. It records the microphone at the right sample rate, plays the agent's voice without gaps, and stops playback when the user interrupts. You write no audio code.
</Info>

<Steps>
  <Step title="Create an agent">
    Skip this step if you already have an agent. You only need its `agent_id`.

    ```bash theme={"system"}
    curl -X POST https://app.eclatira.com/api/v1/agents \
      -H "Authorization: Bearer $ECLATIRA_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "name": "Browser voice demo",
        "description": "Answers questions out loud in a browser tab.",
        "instruction": "You are a concise, friendly assistant. Keep spoken answers under three sentences unless the user asks for detail.",
        "first_message": "Hi, I can hear you. What would you like to talk about?"
      }'
    ```

    The response includes the `agent_id`:

    ```json theme={"system"}
    {
      "agent_id": "9f8c1e4a-2b7d-4c51-9f10-6d3b8a21c7e5",
      "name": "Browser voice demo",
      "realtime": { "enable_webcam": true, "enable_screen_share": true },
      "created_at": "2026-01-14T09:12:44.518000+00:00"
    }
    ```

    Camera and screen share are on by default, so the same agent also works for [camera](/realtime/camera) and [screen](/realtime/screen-share) sessions.
  </Step>

  <Step title="Add a token endpoint to your server">
    This is the only server code you need. It keeps the API key, creates one session, and returns the response to the browser.

    <CodeGroup>
      ```js token-server.mjs (Node) theme={"system"}
      /**
       * Minimal token-minting backend.
       *
       *   ECLATIRA_API_KEY=ek_... ECLATIRA_AGENT_ID=your-agent-id \
       *   ALLOWED_ORIGIN=http://localhost:3000 node token-server.mjs
       */

      import { createServer } from "node:http";
      import { randomUUID } from "node:crypto";

      const API_BASE = process.env.ECLATIRA_API_BASE ?? "https://app.eclatira.com";
      const API_KEY = process.env.ECLATIRA_API_KEY;
      const AGENT_ID = process.env.ECLATIRA_AGENT_ID;
      const PORT = Number(process.env.PORT ?? 8787);

      if (!API_KEY || !AGENT_ID) {
        console.error("Set ECLATIRA_API_KEY and ECLATIRA_AGENT_ID.");
        process.exit(1);
      }

      const ALLOWED_SOURCES = new Set(["audio", "camera", "screen", "text"]);

      const server = createServer(async (request, response) => {
        const url = new URL(request.url, `http://${request.headers.host}`);

        // Your own app's exact origin. Do not reflect the caller's Origin header
        // back, and do not use "*": the browser rejects a wildcard on a request
        // that carries credentials.
        response.setHeader("Access-Control-Allow-Origin", process.env.ALLOWED_ORIGIN ?? "http://localhost:3000");
        response.setHeader("Access-Control-Allow-Credentials", "true");

        if (request.method === "OPTIONS") {
          response.writeHead(204).end();
          return;
        }
        if (url.pathname !== "/api/eclatira/token") {
          response.writeHead(404).end();
          return;
        }

        // Authenticate and authorise *your* user here, before minting. Anyone who
        // can reach this endpoint can start a session that bills your account.

        const source = url.searchParams.get("source") ?? "audio";
        if (!ALLOWED_SOURCES.has(source)) {
          response.writeHead(400, { "content-type": "application/json" });
          response.end(JSON.stringify({ error: `Unsupported source "${source}".` }));
          return;
        }

        try {
          const minted = await fetch(`${API_BASE}/api/v1/realtime/sessions`, {
            method: "POST",
            headers: {
              authorization: `Bearer ${API_KEY}`,
              "content-type": "application/json",
            },
            body: JSON.stringify({
              agent_id: AGENT_ID,
              // Must be unique per concurrent session: this value is the session's
              // identity on the server, so two live sessions sharing one collide.
              user_id: `web_${randomUUID()}`,
              source,
            }),
          });

          const body = await minted.text();
          if (!minted.ok) {
            console.error("Mint failed", minted.status, body);
            response.writeHead(minted.status, { "content-type": "application/json" });
            response.end(body);
            return;
          }

          // Single-use, 60-second token: hand it straight to the browser rather
          // than caching it.
          response.writeHead(200, { "content-type": "application/json", "cache-control": "no-store" });
          response.end(body);
        } catch (error) {
          console.error(error);
          response.writeHead(502, { "content-type": "application/json" });
          response.end(JSON.stringify({ error: "Could not reach Eclatira." }));
        }
      });

      server.listen(PORT, () => console.log(`Token server on http://localhost:${PORT}`));
      ```

      ```ts app/api/eclatira/token/route.ts (Next.js) theme={"system"}
      import { randomUUID } from "node:crypto";
      import { NextResponse } from "next/server";

      const API_BASE = process.env.ECLATIRA_API_BASE ?? "https://app.eclatira.com";
      const ALLOWED_SOURCES = new Set(["audio", "camera", "screen", "text"]);

      // A minted token must never be cached or prerendered.
      export const dynamic = "force-dynamic";

      export async function GET(request: Request) {
        const apiKey = process.env.ECLATIRA_API_KEY;
        const agentId = process.env.ECLATIRA_AGENT_ID;
        if (!apiKey || !agentId) {
          return NextResponse.json(
            { error: "Set ECLATIRA_API_KEY and ECLATIRA_AGENT_ID." },
            { status: 500 },
          );
        }

        // Authenticate and authorise *your* user here, before minting. Anyone who
        // can reach this route can start a session that bills your account.
        // const user = await getSessionUser(); if (!user) return new NextResponse(null, { status: 401 });

        const source = new URL(request.url).searchParams.get("source") ?? "audio";
        if (!ALLOWED_SOURCES.has(source)) {
          return NextResponse.json({ error: `Unsupported source "${source}".` }, { status: 400 });
        }

        const minted = await fetch(`${API_BASE}/api/v1/realtime/sessions`, {
          method: "POST",
          headers: { authorization: `Bearer ${apiKey}`, "content-type": "application/json" },
          body: JSON.stringify({
            agent_id: agentId,
            // Unique per concurrent session: this value is the session's identity
            // on the server, so two live sessions sharing one collide.
            user_id: `web_${randomUUID()}`,
            source,
          }),
          cache: "no-store",
        });

        const body = await minted.text();
        return new NextResponse(body, {
          status: minted.status,
          headers: { "content-type": "application/json", "cache-control": "no-store" },
        });
      }
      ```

      ```python main.py (FastAPI) theme={"system"}
      import os
      import uuid

      import httpx
      from fastapi import FastAPI, HTTPException
      from fastapi.responses import JSONResponse

      app = FastAPI()
      API_BASE = os.environ.get("ECLATIRA_API_BASE", "https://app.eclatira.com")
      ALLOWED_SOURCES = {"audio", "camera", "screen", "text"}


      @app.get("/api/eclatira/token")
      async def mint_session(source: str = "audio") -> JSONResponse:
          # Authenticate and authorise your own user here, before minting.
          if source not in ALLOWED_SOURCES:
              raise HTTPException(status_code=400, detail=f"Unsupported source {source!r}.")

          async with httpx.AsyncClient(timeout=10) as client:
              minted = await client.post(
                  f"{API_BASE}/api/v1/realtime/sessions",
                  headers={"Authorization": f"Bearer {os.environ['ECLATIRA_API_KEY']}"},
                  json={
                      "agent_id": os.environ["ECLATIRA_AGENT_ID"],
                      # Unique per concurrent session.
                      "user_id": f"web_{uuid.uuid4()}",
                      "source": source,
                  },
              )
          return JSONResponse(
              minted.json(),
              status_code=minted.status_code,
              headers={"cache-control": "no-store"},
          )
      ```
    </CodeGroup>

    The endpoint returns this to the browser:

    ```json theme={"system"}
    {
      "session_token": "rt_qN7xK2mPv8LtRz4yWcH1aB6dF3sJ0gUe9nXoQ5rTiYk",
      "expires_at": "2026-01-14T09:13:44.518000+00:00",
      "ws_url": "wss://app.eclatira.com/ws/web_5f1b0e8a-3c42-4d7e-9a18-b06f2d5c81ae?session_token=rt_qN7xK2mPv8LtRz4yWcH1aB6dF3sJ0gUe9nXoQ5rTiYk&source=audio"
    }
    ```

    <Warning>
      Put your own login check where the comment says. Anyone who can call this endpoint can start a session that bills your account.
    </Warning>

    Read the [session rules](/realtime/overview#rules-every-integration-must-follow) before you go live. The short version: each token works once, for 60 seconds, and each session needs a new `user_id`.
  </Step>

  <Step title="Install the library">
    `@eclatira/realtime` is one ES module with no dependencies and no build step.

    <Tabs>
      <Tab title="npm">
        ```bash theme={"system"}
        npm install @eclatira/realtime
        ```

        ```js theme={"system"}
        import { EclatiraSession } from "@eclatira/realtime";
        ```
      </Tab>

      <Tab title="CDN">
        No install. Import it straight from the CDN, pinned to a version:

        ```js theme={"system"}
        import { EclatiraSession } from "https://cdn.jsdelivr.net/npm/@eclatira/realtime@0.1.3/eclatira-realtime.js";
        ```
      </Tab>

      <Tab title="Copy the file">
        Copy `packages/realtime-js/eclatira-realtime.js` into your project and import it by path. Use this if your page must make no external requests.
      </Tab>
    </Tabs>
  </Step>

  <Step title="Build the page">
    Save this file. Serve it from `http://localhost` or HTTPS. The microphone does not work on `file://` pages. Set `TOKEN_URL` to your endpoint from step 2.

    ```html voice.html theme={"system"}
    <!doctype html>
    <html lang="en">
      <head>
        <meta charset="utf-8" />
        <meta name="viewport" content="width=device-width, initial-scale=1" />
        <title>Eclatira browser voice</title>
        <style>
          body { margin: 0; padding: 24px 16px; font: 15px/1.5 system-ui, sans-serif; }
          main { max-width: 640px; margin: 0 auto; }
          button { font: inherit; font-weight: 600; padding: 10px 18px; border-radius: 8px; cursor: pointer; }
          button[disabled] { opacity: 0.45; cursor: not-allowed; }
          #status { margin-left: 12px; color: #6b7280; font-size: 13px; }
          #log { margin-top: 20px; border: 1px solid #e5e7eb; border-radius: 10px; padding: 12px; min-height: 180px; }
          .line { padding: 4px 0; border-bottom: 1px solid #f3f4f6; }
          .line:last-child { border-bottom: 0; }
          .who { font-weight: 600; margin-right: 6px; }
          .err { color: #dc2626; }
        </style>
      </head>
      <body>
        <main>
          <h1>Talk to the agent</h1>
          <button id="start">Start</button>
          <button id="stop" disabled>Stop</button>
          <span id="status">idle</span>
          <div id="log"></div>
        </main>

        <script type="module">
          // Loaded straight from the CDN mirror of the published npm package, pinned
          // to an exact version. Using a bundler? npm install @eclatira/realtime and
          // import "@eclatira/realtime" instead.
          import { EclatiraSession } from "https://cdn.jsdelivr.net/npm/@eclatira/realtime@0.1.3/eclatira-realtime.js";

          const TOKEN_URL = "http://localhost:8787/api/eclatira/token";

          const startButton = document.getElementById("start");
          const stopButton = document.getElementById("stop");
          const statusEl = document.getElementById("status");
          const logEl = document.getElementById("log");

          let session = null;
          let liveLine = null;
          let liveSpeaker = null;

          function write(speaker, text, isError) {
            // Transcripts stream in as revisions of the current utterance, so
            // replace the live line rather than appending a fragment per update.
            if (speaker !== liveSpeaker || !liveLine) {
              liveLine = document.createElement("div");
              liveLine.className = "line";
              liveLine.innerHTML = '<span class="who"></span><span></span>';
              liveLine.querySelector(".who").textContent = speaker;
              logEl.appendChild(liveLine);
              liveSpeaker = speaker;
            }
            const body = liveLine.querySelector("span:last-child");
            body.textContent = text;
            if (isError) body.classList.add("err");
          }

          startButton.addEventListener("click", async () => {
            startButton.disabled = true;
            statusEl.textContent = "connecting";

            session = new EclatiraSession({ tokenUrl: TOKEN_URL });

            session.on("sessioncreated", ({ session_id }) => {
              // The only positive signal that the session is real. `start()`
              // resolves off this message, not off the socket opening.
              console.log("session", session_id);
            });
            session.on("transcript", ({ speaker, text, final }) => {
              write(speaker, text);
              if (final) liveSpeaker = null;
            });
            session.on("speakingchange", ({ user, agent }) => {
              statusEl.textContent = agent ? "agent speaking" : user ? "listening" : "live";
            });
            session.on("sessionended", ({ reason }) => write("system", `session ended: ${reason}`));
            session.on("billing", (message) => write("system", `billing: ${message.type}`));
            session.on("error", (error) => write("error", `${error.code}: ${error.message}`, true));
            session.on("close", ({ code, reason }) => {
              statusEl.textContent = `closed (${code}${reason ? ` ${reason}` : ""})`;
            });
            session.on("stopped", () => {
              startButton.disabled = false;
              stopButton.disabled = true;
            });

            try {
              await session.start({ source: "audio" });
              statusEl.textContent = "live";
              stopButton.disabled = false;
            } catch (error) {
              write("error", `${error.code ?? "start_failed"}: ${error.message}`, true);
              statusEl.textContent = "failed";
              startButton.disabled = false;
            }
          });

          stopButton.addEventListener("click", () => session?.stop());

          // Release the microphone and the concurrency slot if the tab goes away.
          window.addEventListener("beforeunload", () => session?.stop("page_unload"));
        </script>
      </body>
    </html>
    ```

    Open the page and click **Start**. Allow the microphone and start talking.

    <Note>
      Call `session.start()` from a click or other user action. Browsers only allow the microphone and audio playback after one.
    </Note>

    <Accordion title="CORS error when the token endpoint is on another port">
      If your page and your token endpoint are on different origins, set `ALLOWED_ORIGIN` to the page's exact origin. For example, the page is on `:3000` and the Node server is on `:8787`.

      The library sends the request with `credentials: "include"`. Browsers reject `Access-Control-Allow-Origin: *` on these requests. The symptom is a CORS error and `mint_failed` from `session.start()`. A same-origin route, like the Next.js example, avoids this.
    </Accordion>
  </Step>
</Steps>

## Know when the session is live

The WebSocket opening does **not** mean the session started. When the server refuses a session, it accepts the socket first and then closes it with a code. So `onopen` fires even for an expired token or an unpaid account.

The real signal is the first `session_created` message. `session.start()` waits for it:

* It **resolves** when `session_created` arrives.
* It **rejects** with an `EclatiraError` if the socket closes first. The error's `code` names the reason: `invalid_token`, `rate_limited`, `limit_reached`, `billing_blocked`, `source_not_enabled` or `server_error`.

If you use the raw protocol, do the same. Only show "connected" after `session_created`.

## Library reference

### Constructor options

```js theme={"system"}
const session = new EclatiraSession({ tokenUrl: "/api/eclatira/token" });
```

<ResponseField name="tokenUrl" type="string">
  Your token endpoint. It must return `{ ws_url }`. The library adds `?source=...` to the URL. Required unless you pass `getSession`.
</ResponseField>

<ResponseField name="getSession" type="async function">
  Use this instead of `tokenUrl` when a plain GET is not enough. For example, when you need a POST body or your own auth header. It must return `{ ws_url }`.
</ResponseField>

<ResponseField name="frameIntervalMs" type="number" default="2000">
  How often to send an image, in milliseconds. Camera and screen sessions only.
</ResponseField>

<ResponseField name="frameQuality" type="number" default="0.8">
  JPEG quality of each image. Camera and screen sessions only.
</ResponseField>

<ResponseField name="frameMaxEdge" type="number" default="1280">
  Longest side of each image, in pixels. Camera and screen sessions only.
</ResponseField>

### Methods

| Method | What it does |
| - | - |
| `start({ source, videoElement })` | Asks for the microphone (and camera or screen), creates the session and connects. `source` is `audio`, `camera`, `screen` or `text`. `videoElement` shows the local preview. |
| `stop(reason?)` | Ends the session. Closes the socket, stops every camera and microphone track, and closes the audio. |
| `sendText(text)` | Sends a typed message. The agent replies right away, without waiting for speech. |
| `on(event, handler)` | Listens for an event. Returns the session, so you can chain calls. |

### Events

| Event | Payload | Fires when |
| - | - | - |
| `open` | `{}` | The WebSocket opens |
| `sessioncreated` | `{ type, session_id, resumed }` | The server confirms the session. Fires once. |
| `ready` | `{ source }` | Media, socket and audio are all running |
| `transcript` | `{ speaker, text, final }` | Either side's speech is transcribed |
| `agentaudio` | `{ bytes }` | A chunk of agent audio is queued to play |
| `speakingchange` | `{ user, agent }` | The user or the agent starts or stops speaking |
| `toolcall` | `{ type, data }` | The agent calls a tool, or a tool returns |
| `billing` | Server message | `credits_low`, `limit_reached` or `topup_succeeded` |
| `sessionended` | `{ type, reason }` | The agent ended the conversation |
| `error` | `EclatiraError` | Something went wrong on the client, or the server sent an `error` |
| `close` | `{ code, reason }` | The socket closed |
| `stopped` | `{ reason }` | Cleanup finished and devices are released |

<Tip>
  `agentaudio` gives you the size of each audio chunk in bytes. The agent's audio is 24 kHz, 16-bit, mono. So `bytes / 48000` is the chunk's length in seconds.
</Tip>

## Transcripts

`transcript` events send the full text so far for the current turn, not just the new words. Replace the current line each time. Start a new line when `final` is `true`. The example page does this in `write()`.

To keep a record, send each final transcript to your server as it arrives. Or register a [webhook](/webhooks). It gets the full transcript when the session ends.

## End the session

Call `session.stop()` from your stop button. Also call it on `beforeunload`, as the example does.

Your plan limits how many sessions can run at the same time. A session releases its slot when it disconnects. If you reload a page that never called `stop()`, old sessions can hold slots for a while. New connections then close with code `4029`.

<Tip>
  If you keep getting `4029` while you develop, wait for old sessions to drop. Do not keep retrying. Each attempt also uses one of your 60 sessions per hour.
</Tip>

The server has no idle timeout and no maximum session length that you need to handle. You do not need to send a keepalive or ping. In a very long call, the agent keeps its instructions but forgets the oldest part of the conversation.

## Build it without the library

If you cannot use the library, you must build the audio pipeline yourself. Five things must be right. Four of them fail with no error:

1. Record audio at **16 kHz, 16-bit signed little-endian, mono**, with no WAV header. Base64 the raw bytes.
2. Play the agent's audio through a **second** `AudioContext` at 24 kHz.
3. Send audio all the time, including silence. There is no end-of-turn message.
4. Detect the user's speech locally, and **clear the playback queue** the moment they start talking.
5. Send JSON text frames only. Binary frames are rejected.

[Media format](/realtime/media-format) explains each rule in detail. Here is a complete page that follows them:

<Accordion title="voice-raw.html: a full voice client without the library">
  ```html voice-raw.html theme={"system"}
  <!doctype html>
  <html lang="en">
    <head>
      <meta charset="utf-8" />
      <meta name="viewport" content="width=device-width, initial-scale=1" />
      <title>Eclatira browser voice (raw)</title>
      <style>
        body { margin: 0; padding: 24px 16px; font: 15px/1.5 system-ui, sans-serif; }
        main { max-width: 640px; margin: 0 auto; }
        button { font: inherit; font-weight: 600; padding: 10px 18px; border-radius: 8px; cursor: pointer; }
        button[disabled] { opacity: 0.45; cursor: not-allowed; }
        pre { margin-top: 20px; border: 1px solid #e5e7eb; border-radius: 10px; padding: 12px;
              min-height: 200px; white-space: pre-wrap; }
      </style>
    </head>
    <body>
      <main>
        <h1>Talk to the agent (no library)</h1>
        <button id="start">Start</button>
        <button id="stop" disabled>Stop</button>
        <pre id="log"></pre>
      </main>

      <script type="module">
        const TOKEN_URL = "http://localhost:8787/api/eclatira/token";
        const INPUT_SAMPLE_RATE = 16000;
        const OUTPUT_SAMPLE_RATE = 24000;

        // ---------------------------------------------------------------- worklets

        const RECORDER_SOURCE = `
  class EclatiraRecorder extends AudioWorkletProcessor {
    constructor(options) {
      super();
      const opts = (options && options.processorOptions) || {};
      this.targetRate = opts.targetRate || 16000;
      // sampleRate is a global inside an AudioWorkletGlobalScope: the rate the
      // context actually got, which is not always the rate we asked for.
      this.ratio = sampleRate / this.targetRate;
      this.cursor = 0;
      this.tail = 0;

      // Local voice activity detection, used only to cut off playback the moment
      // the user starts talking. The agent runs its own VAD; this one exists
      // because waiting for the server to tell us means the agent talks over the
      // user for the length of the network round trip plus whatever is buffered.
      this.energyThreshold = opts.energyThreshold || 0.015;
      this.speechSeconds = opts.speechSeconds || 0.6;
      this.silenceSeconds = opts.silenceSeconds || 0.8;
      this.speechElapsed = 0;
      this.silenceElapsed = 0;
      this.speaking = false;
    }

    process(inputs) {
      const input = inputs[0] && inputs[0][0];
      if (!input) return true;

      let sum = 0;
      for (let i = 0; i < input.length; i++) sum += input[i] * input[i];
      const energy = Math.sqrt(sum / input.length);
      const frameSeconds = input.length / sampleRate;

      if (energy > this.energyThreshold) {
        this.speechElapsed += frameSeconds;
        this.silenceElapsed = 0;
        if (!this.speaking && this.speechElapsed > this.speechSeconds) {
          this.speaking = true;
          this.port.postMessage({ type: "speech_start" });
        }
      } else {
        this.silenceElapsed += frameSeconds;
        if (this.speaking && this.silenceElapsed > this.silenceSeconds) {
          this.speaking = false;
          this.speechElapsed = 0;
          this.port.postMessage({ type: "speech_end" });
        }
      }

      // Resample to the target rate by linear interpolation. When the context
      // already runs at the target rate this reduces to a straight copy, so the
      // fast path costs nothing; when it does not, this is the difference
      // between a working agent and one that never answers.
      const out = [];
      let position = this.cursor;
      while (position < input.length) {
        const index = Math.floor(position);
        const fraction = position - index;
        const current = index === 0 ? this.tail : input[index - 1];
        const next = input[index];
        const sample = current + (next - current) * fraction;
        const clamped = Math.max(-1, Math.min(1, sample));
        out.push(clamped < 0 ? clamped * 0x8000 : clamped * 0x7fff);
        position += this.ratio;
      }
      this.cursor = position - input.length;
      this.tail = input[input.length - 1];

      if (out.length) {
        const pcm = new Int16Array(out);
        this.port.postMessage({ type: "audio", buffer: pcm.buffer }, [pcm.buffer]);
      }
      return true;
    }
  }
  registerProcessor("eclatira-recorder", EclatiraRecorder);
  `;

        const PLAYER_SOURCE = `
  class EclatiraPlayer extends AudioWorkletProcessor {
    constructor(options) {
      super();
      const opts = (options && options.processorOptions) || {};
      this.sourceRate = opts.sourceRate || 24000;
      this.ratio = this.sourceRate / sampleRate;
      this.queue = [];
      this.chunk = null;
      this.position = 0;
      this.draining = false;

      this.port.onmessage = (event) => {
        const data = event.data;
        if (data.type === "audio") {
          this.queue.push(new Int16Array(data.buffer));
          if (!this.draining) {
            this.draining = true;
            this.port.postMessage({ type: "playback_start" });
          }
        } else if (data.type === "flush") {
          // Barge-in. Everything still queued belongs to a reply the user has
          // already talked over, so it must be dropped rather than played out.
          this.queue = [];
          this.chunk = null;
          this.position = 0;
          this.draining = false;
        }
      };
    }

    nextSample() {
      while (!this.chunk || this.position >= this.chunk.length) {
        if (!this.queue.length) return null;
        this.chunk = this.queue.shift();
        this.position = 0;
      }
      return this.chunk[Math.floor(this.position)];
    }

    process(inputs, outputs) {
      const channel = outputs[0][0];
      for (let i = 0; i < channel.length; i++) {
        const sample = this.nextSample();
        if (sample === null) {
          channel[i] = 0;
          if (this.draining) {
            this.draining = false;
            this.port.postMessage({ type: "playback_end" });
          }
          continue;
        }
        channel[i] = sample / 32768;
        this.position += this.ratio;
      }
      return true;
    }
  }
  registerProcessor("eclatira-player", EclatiraPlayer);
  `;

        // ---------------------------------------------------------------- helpers

        const workletUrls = [];

        function workletUrl(source) {
          const url = URL.createObjectURL(new Blob([source], { type: "application/javascript" }));
          workletUrls.push(url);
          return url;
        }

        function toBase64(arrayBuffer) {
          const bytes = new Uint8Array(arrayBuffer);
          let binary = "";
          const CHUNK = 0x8000; // String.fromCharCode overflows the stack on large inputs.
          for (let i = 0; i < bytes.length; i += CHUNK) {
            binary += String.fromCharCode.apply(null, bytes.subarray(i, i + CHUNK));
          }
          return btoa(binary);
        }

        function fromBase64(value) {
          const binary = atob(value);
          const bytes = new Uint8Array(binary.length);
          for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i);
          return bytes.buffer;
        }

        const logEl = document.getElementById("log");
        function log(line) {
          logEl.textContent += line + "\n";
        }

        // ---------------------------------------------------------------- session

        const startButton = document.getElementById("start");
        const stopButton = document.getElementById("stop");

        let ws = null;
        let stream = null;
        let recorderContext = null;
        let playerContext = null;
        let recorderNode = null;
        let playerNode = null;

        function send(payload) {
          if (ws && ws.readyState === WebSocket.OPEN) ws.send(JSON.stringify(payload));
        }

        function flushPlayback() {
          if (playerNode) playerNode.port.postMessage({ type: "flush" });
        }

        function handleMessage(raw) {
          let message;
          try {
            message = JSON.parse(raw);
          } catch {
            return;
          }

          // Typed messages carry "type"; voice turn and content events do not.
          if (message.type === "session_created") {
            // The only positive signal. Everything before this is unconfirmed:
            // a refused session opens the socket and is closed straight after.
            log(`session ${message.session_id}`);
            return;
          }
          if (message.type === "session_ended") {
            log(`session ended: ${message.reason}`);
            return;
          }
          if (message.type === "credits_low") {
            log(`credits low: ${message.pct_used}% of plan used`);
            return;
          }
          if (message.type === "error") {
            log(`server error [${message.code}] ${message.error}`);
            return;
          }

          // The user talked over the agent: drop the buffered reply.
          if (message.interrupted) flushPlayback();

          for (const part of message.parts || []) {
            if (part.type === "audio/pcm") {
              playerNode.port.postMessage({ type: "audio", buffer: fromBase64(part.data) });
            } else if (part.type === "function_call") {
              log(`tool call: ${part.data.name} ${JSON.stringify(part.data.args)}`);
            } else if (part.type === "function_response") {
              log(`tool result: ${part.data.name}`);
            }
          }
          if (message.input_transcription) {
            log(`user: ${message.input_transcription.text}`);
          }
          if (message.output_transcription) {
            log(`agent: ${message.output_transcription.text}`);
          }
        }

        async function start() {
          // 1. Microphone first: the permission prompt belongs to this click.
          stream = await navigator.mediaDevices.getUserMedia({ audio: true, video: false });

          // 2. Mint through your own backend. One token, one connection, 60s.
          const response = await fetch(`${TOKEN_URL}?source=audio`, { credentials: "include" });
          if (!response.ok) throw new Error(`Token endpoint returned ${response.status}`);
          const session = await response.json();
          if (!session.ws_url) throw new Error("Token endpoint did not return a ws_url.");

          // 3. Playback graph: a second context, at the agent's 24 kHz.
          playerContext = new AudioContext({ sampleRate: OUTPUT_SAMPLE_RATE });
          if (playerContext.state === "suspended") await playerContext.resume();
          await playerContext.audioWorklet.addModule(workletUrl(PLAYER_SOURCE));
          playerNode = new AudioWorkletNode(playerContext, "eclatira-player", {
            processorOptions: { sourceRate: OUTPUT_SAMPLE_RATE },
          });
          playerNode.connect(playerContext.destination);

          // 4. Capture graph, and the check that saves the session.
          recorderContext = new AudioContext({ sampleRate: INPUT_SAMPLE_RATE });
          if (recorderContext.state === "suspended") await recorderContext.resume();
          if (recorderContext.sampleRate !== INPUT_SAMPLE_RATE) {
            log(
              `browser ignored the 16000 Hz request and gave ${recorderContext.sampleRate} Hz; ` +
                "the worklet is resampling"
            );
          }
          await recorderContext.audioWorklet.addModule(workletUrl(RECORDER_SOURCE));
          recorderNode = new AudioWorkletNode(recorderContext, "eclatira-recorder", {
            processorOptions: { targetRate: INPUT_SAMPLE_RATE },
          });

          // 5. Socket. Open ws_url exactly as returned.
          ws = new WebSocket(session.ws_url);
          ws.onopen = () => log("socket open");
          ws.onmessage = (event) => handleMessage(event.data);
          ws.onerror = () => log("socket error");
          ws.onclose = (event) => {
            log(`socket closed: ${event.code}${event.reason ? ` ${event.reason}` : ""}`);
            teardown();
          };

          recorderNode.port.onmessage = ({ data }) => {
            if (data.type === "audio") {
              send({ mime_type: "audio/pcm", data: toBase64(data.buffer) });
            } else if (data.type === "speech_start") {
              // Barge-in: kill the queued reply before the user finishes a word.
              flushPlayback();
            }
          };

          recorderContext.createMediaStreamSource(stream).connect(recorderNode);
        }

        function teardown() {
          if (ws && ws.readyState <= WebSocket.OPEN) {
            ws.onclose = null;
            try {
              ws.close(1000, "client_stopped");
            } catch {
              /* already closing */
            }
          }
          ws = null;

          if (stream) stream.getTracks().forEach((track) => track.stop());
          stream = null;

          for (const context of [recorderContext, playerContext]) {
            if (context && context.state !== "closed") context.close().catch(() => {});
          }
          recorderContext = null;
          playerContext = null;
          recorderNode = null;
          playerNode = null;

          workletUrls.splice(0).forEach(URL.revokeObjectURL);

          startButton.disabled = false;
          stopButton.disabled = true;
        }

        startButton.addEventListener("click", async () => {
          startButton.disabled = true;
          try {
            await start();
            stopButton.disabled = false;
          } catch (error) {
            log(`start failed: ${error.message}`);
            teardown();
          }
        });

        stopButton.addEventListener("click", teardown);
        window.addEventListener("beforeunload", teardown);
      </script>
    </body>
  </html>
  ```

  **Notes**

  * The recorder node is not connected to `destination`. Connecting it would play the user's microphone back through the speakers.
  * There is no required chunk size. This code sends one message per 128-sample block, about every 8 ms at 16 kHz. Messages of 20 to 100 ms also work and cost less overhead.
  * The local speech detector only clears playback. It does not decide turns. The agent does that from the audio.
  * The detector needs 0.6 seconds of speech before it fires. So the agent keeps talking for about 600 ms after the user starts. Shorter sounds, like a cough, do not interrupt the agent. You can lower `speechSeconds` for faster interruptions, but background noise will then interrupt the agent more often.
  * The server also sends `interrupted: true`. It arrives later than the local detector, so use it as a backup.
  * Agent audio chunks are continuous. Play them back to back, with no gaps.
</Accordion>

## Next steps

<CardGroup cols={2}>
  <Card title="Camera" icon="video" href="/realtime/camera">
    Let the agent see the user's webcam.
  </Card>

  <Card title="Screen share" icon="display" href="/realtime/screen-share">
    Let the agent see the user's screen.
  </Card>

  <Card title="WebSocket protocol" icon="plug" href="/realtime-protocol">
    Every message on the socket, field by field.
  </Card>

  <Card title="Errors and close codes" icon="triangle-exclamation" href="/realtime/errors">
    What each code means and how to fix it.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.