> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshooting

> Fix a web session that connects but never speaks, sounds wrong, or keeps disconnecting.

Most broken integrations fail the same way. The WebSocket opens, `session_created` arrives, audio goes out, and nothing comes back. There is no error.

This happens because the server cannot check your audio. Raw audio has no header. So 48 kHz audio looks the same as 16 kHz audio. The server accepts it, and the agent never hears anything it understands.

<Tip>
  **Try this first.** Send one text message on the session:

  ```js theme={"system"}
  ws.send(JSON.stringify({ mime_type: "text/plain", data: "Say the word banana." }));
  ```

  * **The agent replies?** Your token, socket, agent and billing all work. The problem is in your audio code.
  * **No reply?** The problem is not your audio. Check the steps below from the top.
</Tip>

## Add logging first

Attach this helper right after `new WebSocket(...)`. It logs `session_created`, every error, the close code, and how many bytes of audio you send per second.

```js theme={"system"}
/**
 * Drop-in diagnostics for an Eclatira realtime WebSocket.
 *
 * Wraps send() to measure what you are actually transmitting, and logs every
 * typed server message. Returns a function that stops the reporting timer.
 */
export function attachRealtimeDiagnostics(ws, { intervalMs = 1000 } = {}) {
  const stats = { pcmBytes: 0, pcmMessages: 0, frames: 0, texts: 0, unknown: 0 };
  let sessionId = null;
  let firstError = null;
  let sawSessionCreated = false;

  const nativeSend = ws.send.bind(ws);
  ws.send = (payload) => {
    if (typeof payload !== "string") {
      console.error(
        "[eclatira] Outbound binary frame. The server only reads text frames " +
          "and will answer binary_frame_unsupported."
      );
      return nativeSend(payload);
    }
    try {
      const message = JSON.parse(payload);
      const type = String(message.mime_type || "").split(";")[0].trim().toLowerCase();
      if (type === "audio/pcm" || type === "audio/l16") {
        const b64 = message.data || "";
        const padding = b64.endsWith("==") ? 2 : b64.endsWith("=") ? 1 : 0;
        stats.pcmBytes += Math.floor((b64.length * 3) / 4) - padding;
        stats.pcmMessages += 1;
      } else if (type === "image/jpeg" || type === "image/jpg") {
        stats.frames += 1;
      } else if (type === "text/plain" || type === "text") {
        stats.texts += 1;
      } else {
        stats.unknown += 1;
        console.warn("[eclatira] Outbound mime_type the server does not accept:", message.mime_type);
      }
    } catch {
      console.error("[eclatira] Outbound payload is not JSON. The server will answer invalid_json.");
    }
    return nativeSend(payload);
  };

  ws.addEventListener("message", (event) => {
    let message;
    try {
      message = JSON.parse(event.data);
    } catch {
      return;
    }
    if (message.type === "session_created") {
      sawSessionCreated = true;
      sessionId = message.session_id;
      console.log("[eclatira] session_created", sessionId, "resumed:", message.resumed);
    } else if (message.type === "error") {
      if (!firstError) firstError = message;
      console.error("[eclatira] server error:", message.code, "-", message.error);
    } else if (message.type) {
      console.log("[eclatira]", message.type, message);
    } else if (message.input_transcription) {
      console.log("[eclatira] the model heard:", message.input_transcription.text);
    }
  });

  ws.addEventListener("close", (event) => {
    clearInterval(timer);
    console.log("[eclatira] closed", {
      code: event.code,
      reason: event.reason,
      sessionId,
      sawSessionCreated,
      firstError,
    });
  });

  const timer = setInterval(() => {
    const seconds = intervalMs / 1000;
    console.log(
      `[eclatira] out: ${(stats.pcmBytes / seconds / 1000).toFixed(1)} kB/s PCM ` +
        `over ${stats.pcmMessages} messages, ${stats.frames} frames, ${stats.texts} text`
    );
    stats.pcmBytes = 0;
    stats.pcmMessages = 0;
    stats.frames = 0;
    stats.texts = 0;
    stats.unknown = 0;
  }, intervalMs);

  return () => clearInterval(timer);
}
```

A correct voice session logs **`32.0 kB/s PCM`** while the microphone is on. Any other number points to the bug.

## The agent never replies

Go through these checks in order.

<Steps>
  <Step title="Did session_created arrive?">
    The server sends `session_created` first, once, on every session.

    * **It never arrived.** The server closed the socket during setup. Read `event.code` in `onclose` and look it up in [Close codes](/realtime/errors#close-codes). Do not retry with the same token. It is already used, so you will get `4001` and hide the real cause.
    * **It arrived.** Go to the next step.
  </Step>

  <Step title="Did you get any error messages?">
    The server sends `{"type": "error", "code": "..."}` for anything it cannot use. The session stays open.

    * **`no_audio_received`**: you are sending images but no audio. Go to step 6.
    * **Any other code**: look it up in [Error messages](/realtime/errors#error-messages).
    * **No errors**: the server accepts your messages. The problem is inside the audio itself. Go to step 3.

    <Note>
      `text` sessions do not send these errors. They ignore audio and images without saying anything.
    </Note>
  </Step>

  <Step title="Is audio being sent?">
    Check the byte rate in the logging helper. 16,000 samples × 2 bytes × 1 channel = **32,000 bytes per second**.

    | Rate | Meaning |
    | - | - |
    | `32.0 kB/s` | Correct. |
    | `0.0 kB/s` | Nothing is sent. The audio graph is not connected, the `AudioContext` is suspended, or the microphone is muted. |
    | `96.0 kB/s` | You record at 48 kHz and do not resample. Go to step 4. |
    | `88.2 kB/s` | You record at 44.1 kHz and do not resample. Go to step 4. |
    | `64.0 kB/s` | Stereo, or Float32 samples. Go to step 5. |
    | `192.0 kB/s` | Float32 at 48 kHz. Both problems. Steps 4 and 5. |

    An `AudioContext` created outside a click starts suspended. It sends nothing until `resume()` finishes, with no error.

    Send audio all the time, including silence. The agent uses the silence to tell when the user has stopped talking.
  </Step>

  <Step title="Is the sample rate right?">
    This is the most common cause. `new AudioContext({ sampleRate: 16000 })` is only a request. Safari and Firefox often ignore it and use 44.1 or 48 kHz. Nothing tells you.

    Check what you actually got:

    ```js theme={"system"}
    /**
     * Report the rate the browser actually gave you, which is not always the rate
     * you asked for. Run this once, on the device and browser you are debugging.
     */
    async function reportAudioContextRate() {
      const context = new AudioContext({ sampleRate: 16000 });
      try {
        if (context.state === "suspended") await context.resume();
        console.log("requested 16000, got", context.sampleRate);
        if (context.sampleRate !== 16000) {
          console.warn(
            "This browser refused the requested rate. You must resample to 16000 " +
              "before sending, or the agent will never respond."
          );
        }
        return context.sampleRate;
      } finally {
        await context.close();
      }
    }

    reportAudioContextRate();
    ```

    If the rate is wrong, the agent hears your voice three times too slow and far too deep. It never detects speech. There is no error anywhere.

    Two fixes: use the [client library](/realtime/browser-voice), which resamples for you, or resample yourself:

    ```js theme={"system"}
    /**
     * Resample mono Float32 audio to a target rate by linear interpolation and
     * convert to the 16-bit signed little-endian PCM the agent accepts.
     *
     * `tail` carries the last sample of the previous block so block boundaries do
     * not click. Pass the returned `tail` and `cursor` back in on the next call.
     *
     * @param {Float32Array} input   One block of captured audio.
     * @param {number} inputRate     The rate the AudioContext actually reported.
     * @param {number} targetRate    16000.
     * @param {{cursor: number, tail: number}} state  Carried between calls.
     * @returns {Int16Array}
     */
    function resampleToInt16(input, inputRate, targetRate, state) {
      const ratio = inputRate / targetRate;
      const out = [];
      let position = state.cursor;
      while (position < input.length) {
        const index = Math.floor(position);
        const fraction = position - index;
        const current = index === 0 ? state.tail : input[index - 1];
        const next = input[index];
        const sample = current + (next - current) * fraction;
        const clamped = Math.max(-1, Math.min(1, sample));
        out.push(clamped < 0 ? clamped * 0x8000 : clamped * 0x7fff);
        position += ratio;
      }
      state.cursor = position - input.length;
      state.tail = input[input.length - 1];
      return new Int16Array(out);
    }

    // Usage, inside an AudioWorkletProcessor or ScriptProcessor callback:
    const resampleState = { cursor: 0, tail: 0 };
    function onAudioBlock(float32Block, contextSampleRate, send) {
      const pcm = resampleToInt16(float32Block, contextSampleRate, 16000, resampleState);
      if (pcm.length) send(pcm.buffer);
    }
    ```

    <Warning>
      Do not resample by dropping samples. It distorts the audio. The agent may then hear the wrong words, which is harder to spot than silence.
    </Warning>
  </Step>

  <Step title="Is the format right?">
    The format is raw PCM: **16 kHz, 16-bit signed little-endian, mono, no WAV header**, base64-encoded. See [Media format](/realtime/media-format#input-audio).

    Four mistakes all look like a silent agent:

    * **A WAV header.** You are probably encoding whole files instead of streaming.
    * **Float32 samples.** The byte rate shows `64.0 kB/s` at 16 kHz.
    * **Big-endian samples.** Only happens if you build bytes with `DataView.setInt16` and leave out `littleEndian`. It defaults to big-endian.
    * **Stereo.** Two channels read as one fast, garbled channel.

    Convert Float32 to Int16 like this:

    ```js theme={"system"}
    /**
     * Convert clamped mono Float32 audio in [-1, 1] to 16-bit signed PCM.
     * The two scale factors are deliberate: Int16 spans -32768..32767, so the
     * negative side is scaled by 0x8000 and the positive side by 0x7fff.
     */
    function floatToPcm16(input) {
      const output = new Int16Array(input.length);
      for (let i = 0; i < input.length; i++) {
        const sample = Math.max(-1, Math.min(1, input[i]));
        output[i] = sample < 0 ? sample * 0x8000 : sample * 0x7fff;
      }
      return output;
    }

    /** Base64-encode an ArrayBuffer without overflowing the call stack. */
    function toBase64(arrayBuffer) {
      const bytes = new Uint8Array(arrayBuffer);
      let binary = "";
      const CHUNK = 0x8000;
      for (let i = 0; i < bytes.length; i += CHUNK) {
        binary += String.fromCharCode.apply(null, bytes.subarray(i, i + CHUNK));
      }
      return btoa(binary);
    }

    /** Send one block of already-resampled 16 kHz PCM. */
    function sendPcm(ws, int16Array) {
      ws.send(JSON.stringify({ mime_type: "audio/pcm", data: toBase64(int16Array.buffer) }));
    }
    ```

    Send small chunks as you record. 20 to 100 ms per message works well. Do not send a whole sentence in one message.
  </Step>

  <Step title="Camera or screen: are you sending microphone audio?">
    The agent only replies to speech. **Images alone never get a reply.** After about 8 seconds of images with no audio, the server sends `no_audio_received`.

    For screen share, the audio from `getDisplayMedia` is the computer's sound, not the microphone. Get the two separately and merge them:

    ```js theme={"system"}
    /**
     * Acquire a screen share plus the microphone as one MediaStream.
     *
     * getDisplayMedia's audio track is system or tab audio, so it is refused here
     * and the microphone is acquired separately. A screen session with no
     * microphone connects and then never responds.
     *
     * @param {() => void} onUserStoppedSharing Called when the browser's own
     *   "Stop sharing" control ends the track. That control does not touch the
     *   WebSocket, so nothing else will tell you.
     */
    async function acquireScreenAndMicrophone(onUserStoppedSharing) {
      const display = await navigator.mediaDevices.getDisplayMedia({
        video: { width: { ideal: 1280 }, height: { ideal: 720 } },
        audio: false,
      });

      let mic;
      try {
        mic = await navigator.mediaDevices.getUserMedia({ audio: true, video: false });
      } catch (error) {
        display.getTracks().forEach((track) => track.stop());
        throw new Error(
          "Screen sharing needs the microphone too: the agent replies to speech, " +
            "so a screen session with no audio will connect and then never respond."
        );
      }

      const [videoTrack] = display.getVideoTracks();
      if (videoTrack) videoTrack.addEventListener("ended", onUserStoppedSharing);

      return {
        stream: new MediaStream([...display.getVideoTracks(), ...mic.getAudioTracks()]),
        release: () => {
          display.getTracks().forEach((track) => track.stop());
          mic.getTracks().forEach((track) => track.stop());
        },
      };
    }
    ```
  </Step>

  <Step title="Does a text message get a reply?">
    A `text/plain` message makes the agent reply right away, with no audio involved. It works on every session type.

    ```js theme={"system"}
    /**
     * Force a turn without audio, to separate a socket/agent fault from an audio
     * pipeline fault. Works on audio, camera and screen sessions.
     */
    function bisectWithText(ws) {
      ws.addEventListener("message", (event) => {
        const message = JSON.parse(event.data);
        if (message.output_transcription) {
          console.log("AGENT REPLIED:", message.output_transcription.text);
        }
        if (message.parts) {
          for (const part of message.parts) {
            if (part.type === "audio/pcm") console.log("agent audio:", part.data.length, "base64 chars");
          }
        }
      });
      ws.send(JSON.stringify({ mime_type: "text/plain", data: "Say the word banana and nothing else." }));
    }
    ```

    Send `text/plain` data as plain text, **not** base64. If you base64 it, the agent reads the code out loud.

    | Result | What it means |
    | - | - |
    | The agent replies | Everything except your audio works. Go back to steps 3 to 5. |
    | No reply, no error | The problem is above the audio. Check that `session_created` arrived. Then check the agent: its instructions, or a tool call that never returns. |
    | An error arrives | Look up the code in [Error messages](/realtime/errors#error-messages). |

    You can also test with no browser at all. Create a session:

    ```bash theme={"system"}
    curl -X POST https://app.eclatira.com/api/v1/realtime/sessions \
      -H "Authorization: Bearer $ECLATIRA_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "agent_id": "YOUR_AGENT_ID",
        "user_id": "debug-session-1",
        "source": "audio"
      }'
    ```

    Then run this script with the `ws_url` within 60 seconds:

    ```python theme={"system"}
    """Prove the socket and agent are healthy, independent of any audio pipeline.

    Usage:
        pip install websockets
        python bisect.py "wss://app.eclatira.com/ws/debug-session-1?session_token=...&source=audio"

    Sends one text/plain message, which forces a turn with no audio, and prints
    every message the server sends for twenty seconds.
    """

    import asyncio
    import json
    import sys

    import websockets


    async def main(ws_url: str) -> None:
        async with websockets.connect(ws_url) as ws:

            async def read() -> None:
                async for raw in ws:
                    message = json.loads(raw)
                    kind = message.get("type")
                    if kind == "error":
                        print("ERROR", message["code"], "-", message["error"])
                    elif kind:
                        print(kind, message)
                        continue
                    for part in message.get("parts") or []:
                        if part["type"] == "audio/pcm":
                            print(f"audio part: {len(part['data'])} base64 chars")
                        else:
                            print(part["type"], part["data"])
                    if message.get("output_transcription"):
                        print("agent said:", message["output_transcription"]["text"])
                    if message.get("input_transcription"):
                        print("agent heard:", message["input_transcription"]["text"])

            reader = asyncio.create_task(read())
            await ws.send(
                json.dumps({"mime_type": "text/plain", "data": "Say the word banana and nothing else."})
            )
            await asyncio.sleep(20)
            reader.cancel()


    asyncio.run(main(sys.argv[1]))
    ```

    If the script gets a reply and your browser does not, the problem is in your browser code. It is not your account, agent or plan.
  </Step>
</Steps>

## Other problems

<AccordionGroup>
  <Accordion title="The agent talks over me when I interrupt">
    Your app must stop the agent's audio. When the user starts talking, **clear the playback queue right away**. Otherwise the agent keeps playing everything already buffered.

    Detect speech locally on the microphone, and clear the queue when speech starts. The reference code uses an energy threshold of `0.015`, 0.6 seconds of speech to start, and 0.8 seconds of silence to stop.

    The server also sends `interrupted: true`. It arrives later, so use it as a backup. See [Barge-in](/realtime/media-format#barge-in).

    With these settings, the agent keeps talking for about 600 ms after the user starts. A very short word, like "wait", may not interrupt it. You can lower `speechSeconds` to react faster. Background noise will then interrupt the agent more often.
  </Accordion>

  <Accordion title="The agent's audio is choppy or clicks">
    The agent's audio is continuous. Gaps come from how you play it.

    * **One `AudioBufferSourceNode` per chunk** adds a gap between chunks. Use one continuous queue instead.
    * **Clicks between chunks** mean chunks are lost or out of order. Play them in the order they arrive.
    * **Clicks all the time** mean clipping. Clamp Float32 samples to `[-1, 1]` before you convert.

    Play audio through a **separate `AudioContext` at 24000 Hz**.
  </Accordion>

  <Accordion title="The agent hears itself and keeps interrupting itself">
    The microphone picks up the agent's voice from the speakers. The agent thinks the user is talking and stops.

    * Turn on echo cancellation: `getUserMedia({ audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true } })`.
    * Test with headphones. If the problem goes away, it is echo.
    * Check `input_transcription`. If it contains the agent's own words, this is the cause.
  </Accordion>

  <Accordion title="The agent sounds too fast, too slow, or garbled">
    Speed and pitch problems always mean a sample rate mismatch.

    * **Agent too fast and high:** you play 24 kHz audio in a faster context without adjusting. Create the playback context at 24000.
    * **Agent too slow and deep:** you play 24 kHz audio as if it were 16 kHz.
    * **Agent mishears you:** your input audio is at the wrong rate or distorted. See steps 4 and 5.
    * **Both wrong:** you use one `AudioContext` for recording and playback. Use two.
  </Accordion>

  <Accordion title="Screen share connects, but the agent says it cannot see">
    Check these in order:

    1. **Wrong `source`.** The default is `audio`. An `audio` session tells the agent it cannot see anything, even if you send images. Create the session with `"source": "screen"` or `"camera"`.
    2. **No microphone audio.** No turn starts. See step 6.
    3. **The screen just changed.** The agent gets at most one image per second. Wait a moment after switching tabs.
    4. **Blank images.** Drawing a video before it has loaded gives a black image. Check `video.videoWidth && video.readyState >= 2` before drawing.
    5. **Images too small.** Use a longest side of 1280 px and quality 0.8.
    6. **The user stopped sharing.** The browser's **Stop sharing** button ends the video track but not the socket. Listen for the track's `ended` event.
  </Accordion>

  <Accordion title="It works on localhost but breaks when deployed">
    There are two separate causes.

    **The browser cannot create sessions.** The API accepts browser requests from `http://localhost:3000`, but not from your domain. So creating sessions from the browser works locally and fails with a CORS error when deployed. Create sessions on your server. See [Authentication](/authentication#keep-your-key-on-the-server).

    **The page is not HTTPS.** The microphone and screen capture only work on secure pages. On a deployed `http://` page, `navigator.mediaDevices` is `undefined`. Serve over HTTPS.
  </Accordion>

  <Accordion title="429 when creating sessions while debugging">
    You can create 60 sessions per hour and 500 per day per workspace. Each connection attempt uses one. These use them up fast:

    * Creating a session when a component mounts. React Strict Mode doubles this in development.
    * Creating a session on every hot reload.
    * Reconnecting in a loop after every close.

    Create sessions on a user action, like a button click. Add a limit and a delay to automatic reconnects.
  </Accordion>

  <Accordion title="start() rejects with invalid_token">
    The socket closed with `4001`. Usually the token was already used, often by a double-run effect, a hot reload, or a retry. It may also have expired after 60 seconds. See [4001: invalid token](/realtime/errors#4001-invalid-token).
  </Accordion>

  <Accordion title="4029 every time I reload the page">
    Read the close reason.

    * **"Too many connections", or `1006` with no reason:** too many attempts from your IP. Wait a minute.
    * **"Too many concurrent sessions for your plan":** old sessions still hold slots. Always call `stop()` on `beforeunload`. A session closed without cleanup can hold its slot for up to 2 hours.

    Use a new `user_id` for each session, and reuse it only for retries of that same session.
  </Accordion>

  <Accordion title="The permission prompt never appears">
    * **Not HTTPS.** `navigator.mediaDevices` does not exist on insecure pages.
    * **No user action.** `getDisplayMedia` must be called directly in a click handler.
    * **Blocked before.** If the user blocked the microphone once, the browser remembers. They must reset it in site settings.
    * **Inside an iframe.** The iframe needs `allow="microphone; camera; display-capture"`.
    * **Permissions-Policy.** A strict `Permissions-Policy` header on your page blocks it.
    * **No device.** With no microphone, you get `NotFoundError`.

    Log the error's `name`. `NotAllowedError`, `NotFoundError`, `InvalidStateError` and `TypeError` each point to a different cause.
  </Accordion>
</AccordionGroup>

## Known limits

These are not bugs in your code.

* **You cannot fetch a web session through `GET /api/v1/calls/{call_id}`.** Collect the transcript in the browser, or use a [webhook](/webhooks).
* **A dropped session cannot be resumed.** Create a new session and connect again.
* **No keepalive is needed.** There is no idle timeout or maximum length to handle. If sessions end after the same time every time, check for a proxy or load balancer timeout on your side.

## Still stuck?

Collect these before you ask for help:

<CardGroup cols={2}>
  <Card title="Close code and reason" icon="plug">
    `event.code` and `event.reason` from `onclose`.
  </Card>

  <Card title="The first error" icon="triangle-exclamation">
    The first `error` message, with its `code` and text.
  </Card>

  <Card title="The real sample rate" icon="wave-square">
    `context.sampleRate`, plus the browser and version.
  </Card>

  <Card title="session_id" icon="fingerprint">
    From `session_created`. It links your session to our logs.
  </Card>
</CardGroup>

Also include:

* The PCM byte rate from the logging helper. `32.0` is correct.
* The `source`, `user_id` and HTTP status of the session request. **Never** the `session_token` or your API key.
* Whether the text message test got a reply.

This helper collects most of it for you:

```js theme={"system"}
/**
 * Collect everything worth reporting about a failed session, in one object.
 * Attach this to the socket alongside the diagnostics helper above.
 */
export function collectSessionReport(ws) {
  const report = {
    sessionId: null,
    resumed: null,
    sawSessionCreated: false,
    firstError: null,
    closeCode: null,
    closeReason: null,
    userAgent: navigator.userAgent,
    requestedSampleRate: 16000,
    actualSampleRate: null,
    secureContext: window.isSecureContext,
  };

  const probe = new AudioContext({ sampleRate: 16000 });
  report.actualSampleRate = probe.sampleRate;
  probe.close();

  ws.addEventListener("message", (event) => {
    let message;
    try {
      message = JSON.parse(event.data);
    } catch {
      return;
    }
    if (message.type === "session_created") {
      report.sawSessionCreated = true;
      report.sessionId = message.session_id;
      report.resumed = message.resumed;
    } else if (message.type === "error" && !report.firstError) {
      report.firstError = { code: message.code, error: message.error };
    }
  });

  ws.addEventListener("close", (event) => {
    report.closeCode = event.code;
    report.closeReason = event.reason;
    console.log("[eclatira] session report\n" + JSON.stringify(report, null, 2));
  });

  return report;
}
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.