> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Media format

> The exact audio and image formats a web session sends and receives, with working code for each.

This page is the byte-level spec for a web session. Use it if you write your own client instead of using the library. If a session connects and the agent never speaks, the cause is almost always on this page.

<Warning>
  **Most format mistakes fail with no error.** The server does not inspect your audio. It passes the bytes to the agent as 16 kHz, 16-bit, mono PCM, whatever they really are. A wrong sample rate, bit depth, byte order, channel count or a WAV header all look the same: the socket stays open, no error arrives, and the agent never replies.
</Warning>

## Summary

| Stream | Rate | Samples | Channels | Container | Sent as |
| - | - | - | - | - | - |
| Audio you send | 16000 Hz | 16-bit signed, little-endian | 1 | None | `{"mime_type": "audio/pcm", "data": "<base64>"}` |
| Audio you receive | 24000 Hz | 16-bit signed, little-endian | 1 | None | A `parts` entry with `type: "audio/pcm"` |
| Images you send | | JPEG | | JPEG | `{"mime_type": "image/jpeg", "data": "<base64>"}` |

The input and output rates are different. You cannot change either one.

The agent only takes a turn when it hears speech in the audio. There is no end-of-turn message. Images never start a turn.

## Input audio

### Required format

| Property | Value | Notes |
| - | - | - |
| Sample rate | `16000` Hz | Not 44100, 48000 or 24000. 24000 is the output rate. |
| Bit depth | 16 bits | Web Audio gives you Float32. You must convert it. |
| Sample type | Signed integer | Range −32768 to 32767. |
| Byte order | Little-endian | `Int16Array` is already little-endian in every browser. |
| Channels | 1 (mono) | Not stereo. |
| Container | None | Raw samples only. No WAV header, no WebM, no Opus, no MP3. |

Four samples with values `0, 1000, -1000, 32767` look like this in bytes:

```text theme={"system"}
00 00   E8 03   18 FC   FF 7F
 └─ 0    └─1000  └─-1000  └─32767
```

The byte count is always `samples × 2`. There is no minimum length and no alignment rule.

### Convert Float32 to Int16

Web Audio gives you samples from −1 to 1. Scale negative and positive values differently, because the Int16 range is not symmetric:

```js theme={"system"}
/**
 * Convert one Float32 sample in [-1, 1] to a 16-bit signed integer.
 * The negative side scales by 0x8000 (32768) and the positive side by
 * 0x7fff (32767), because the range is asymmetric. Using 0x8000 on both
 * sides makes +1.0 overflow to -32768: a full-scale sample becomes a
 * full-scale sample of the opposite sign, which is an audible click on
 * every loud syllable.
 */
function floatToInt16(sample) {
  const clamped = Math.max(-1, Math.min(1, sample));
  return clamped < 0 ? clamped * 0x8000 : clamped * 0x7fff;
}

/** Convert a whole Float32Array to little-endian Int16 bytes. */
function encodePcm16(floats) {
  const out = new Int16Array(floats.length);
  for (let i = 0; i < floats.length; i++) out[i] = floatToInt16(floats[i]);
  return out.buffer; // Int16Array is little-endian on every supported platform.
}
```

Always clamp first. Web Audio can produce values outside −1 to 1. Without the clamp, those values wrap around and cause clicks.

### The AudioContext sample-rate problem

This is the most common reason a correct-looking client never gets a reply.

```js theme={"system"}
const context = new AudioContext({ sampleRate: 16000 });
```

`sampleRate` is a **request**. The browser can give you a different rate, and some do:

| Behavior | Browsers |
| - | - |
| Honors 16000 and resamples the microphone | Chrome and other Chromium browsers, in most setups |
| Often uses the hardware rate (44100 or 48000) | Safari, Firefox |

Nothing warns you. You send 48 kHz audio labeled as 16 kHz. The agent hears one second of speech as three seconds of slow, deep sound. It never detects speech.

Always check the rate you got:

```js theme={"system"}
const context = new AudioContext({ sampleRate: 16000 });
if (context.sampleRate !== 16000) {
  console.warn(
    `AudioContext runs at ${context.sampleRate} Hz, not 16000. ` +
    `Audio must be resampled before sending or the agent will never respond.`
  );
}
```

Do not check the browser name. Read `context.sampleRate` and resample from that. The resampler below costs nothing when the rate already matches.

### The resampling worklet

This is the same recorder the [client library](/realtime/browser-voice) uses. It reads the real rate, resamples to 16 kHz, converts to Int16, and posts the bytes to the main thread. It also detects speech locally for [barge-in](#barge-in).

Save it as `eclatira-recorder.js` and serve it from your own site.

```js theme={"system"}
// eclatira-recorder.js
class EclatiraRecorder extends AudioWorkletProcessor {
  constructor(options) {
    super();
    const opts = (options && options.processorOptions) || {};
    this.targetRate = opts.targetRate || 16000;
    // `sampleRate` is a global inside AudioWorkletGlobalScope: the rate the
    // context actually got, which is not always the rate that was requested.
    this.ratio = sampleRate / this.targetRate;

    // Resampler state that MUST survive between process() calls.
    this.cursor = 0; // fractional read position carried into the next block
    this.tail = 0;   // last input sample of the previous block

    // Local voice activity detection, used only to cut playback off the
    // instant the user starts talking.
    this.energyThreshold = opts.energyThreshold || 0.015;
    this.speechSeconds = opts.speechSeconds || 0.6;
    this.silenceSeconds = opts.silenceSeconds || 0.8;
    this.speechElapsed = 0;
    this.silenceElapsed = 0;
    this.speaking = false;
  }

  process(inputs) {
    const input = inputs[0] && inputs[0][0];
    if (!input) return true;

    let sum = 0;
    for (let i = 0; i < input.length; i++) sum += input[i] * input[i];
    const energy = Math.sqrt(sum / input.length);
    const frameSeconds = input.length / sampleRate;

    if (energy > this.energyThreshold) {
      this.speechElapsed += frameSeconds;
      this.silenceElapsed = 0;
      if (!this.speaking && this.speechElapsed > this.speechSeconds) {
        this.speaking = true;
        this.port.postMessage({ type: "speech_start" });
      }
    } else {
      this.silenceElapsed += frameSeconds;
      if (this.speaking && this.silenceElapsed > this.silenceSeconds) {
        this.speaking = false;
        this.speechElapsed = 0;
        this.port.postMessage({ type: "speech_end" });
      }
    }

    const out = [];
    let position = this.cursor;
    while (position < input.length) {
      const index = Math.floor(position);
      const fraction = position - index;
      const current = index === 0 ? this.tail : input[index - 1];
      const next = input[index];
      const sample = current + (next - current) * fraction;
      const clamped = Math.max(-1, Math.min(1, sample));
      out.push(clamped < 0 ? clamped * 0x8000 : clamped * 0x7fff);
      position += this.ratio;
    }
    // Carry the fractional overshoot and the block's last sample forward.
    this.cursor = position - input.length;
    this.tail = input[input.length - 1];

    if (out.length) {
      const pcm = new Int16Array(out);
      this.port.postMessage({ type: "audio", buffer: pcm.buffer }, [pcm.buffer]);
    }
    return true;
  }
}
registerProcessor("eclatira-recorder", EclatiraRecorder);
```

Wire it up on the main thread:

```js theme={"system"}
/**
 * Capture the microphone, resample to 16 kHz, and send each block on the
 * realtime socket. `ws` is an open WebSocket; `stream` is a MediaStream with
 * at least one audio track.
 */
async function startCapture(ws, stream, { onSpeechStart, onSpeechEnd } = {}) {
  const context = new AudioContext({ sampleRate: 16000 });
  if (context.state === "suspended") await context.resume();
  if (context.sampleRate !== 16000) {
    console.warn(`Context at ${context.sampleRate} Hz; resampling to 16000.`);
  }

  await context.audioWorklet.addModule("/eclatira-recorder.js");
  const node = new AudioWorkletNode(context, "eclatira-recorder", {
    processorOptions: { targetRate: 16000 },
  });

  node.port.onmessage = ({ data }) => {
    if (data.type === "audio") {
      ws.send(JSON.stringify({ mime_type: "audio/pcm", data: toBase64(data.buffer) }));
    } else if (data.type === "speech_start") {
      onSpeechStart?.();
    } else if (data.type === "speech_end") {
      onSpeechEnd?.();
    }
  };

  context.createMediaStreamSource(stream).connect(node);
  return { context, node };
}

/** btoa() on a large binary string overflows the call stack; chunk it. */
function toBase64(arrayBuffer) {
  const bytes = new Uint8Array(arrayBuffer);
  let binary = "";
  const CHUNK = 0x8000;
  for (let i = 0; i < bytes.length; i += CHUNK) {
    binary += String.fromCharCode.apply(null, bytes.subarray(i, i + CHUNK));
  }
  return btoa(binary);
}
```

<Note>
  The library includes this worklet inside its own file and loads it from a `Blob` URL. You do not need to host anything if you use the library.
</Note>

### Why the resampler keeps state

`process()` runs once per block of 128 samples. The resampler steps through the input by a fraction, so two values must carry over from one block to the next:

| State | Why it matters |
| - | - |
| `this.cursor` | The leftover fractional position. Without it, each block emits slightly too many samples. The stream drifts ahead of real time, and the sound jumps at every block. |
| `this.tail` | The last sample of the previous block. Without it, the first sample of each block becomes `NaN`, then `0`. You hear a buzz hundreds of times a second. |

These bugs make a hand-written resampler sound almost right, while speech detection fails.

### How often to send

| Question | Answer |
| - | - |
| Minimum chunk size | None. Only an empty payload is rejected. |
| Maximum chunk size | No app limit. The transport limit is [16 MiB per message](#message-size-limit). |
| Required timing | None. |
| Good range | 20 to 100 ms of audio per message. |
| What the library does | One message per 128-sample block. That is 8 ms at 16 kHz. |
| End-of-turn message | **None.** The agent detects turns from the audio. |

Send audio **all the time, including silence**. The agent needs the silence to know the user stopped talking. If you only send while the user speaks, the agent either interrupts all the time or never answers.

To make the agent reply without speech, send text. This starts a turn right away:

```js theme={"system"}
ws.send(JSON.stringify({ mime_type: "text/plain", data: "What is on my screen?" }));
```

`text/plain` data is **plain UTF-8, not base64**. It is the only type that is not base64.

## Output audio

The agent's voice arrives as `parts` entries:

```json theme={"system"}
{
  "is_partial": true,
  "parts": [{ "type": "audio/pcm", "data": "AAAA6AMY/P9/" }],
  "input_transcription": null,
  "output_transcription": null
}
```

| Property | Value |
| - | - |
| Sample rate | `24000` Hz |
| Bit depth | 16 bits |
| Sample type | Signed integer |
| Byte order | Little-endian |
| Channels | 1 (mono) |
| Container | None |
| Continuity | Chunks are continuous. Play them back to back in arrival order, with no gaps. |

Chunk sizes vary. Do not assume a fixed size.

### You need two AudioContexts

An `AudioContext` has one sample rate. Recording needs 16000 and playback needs 24000. Create two:

```js theme={"system"}
const captureContext = new AudioContext({ sampleRate: 16000 });
const playbackContext = new AudioContext({ sampleRate: 24000 });
```

With one context, one side is always wrong. At 16000, the agent sounds slow and low. At 24000, your microphone audio goes out at the wrong rate and the agent never replies.

### The playback worklet

Playing each chunk with its own `BufferSource` adds gaps. It also gives you nothing to clear on barge-in. Use a queue in a worklet.

Save it as `eclatira-player.js`:

```js theme={"system"}
// eclatira-player.js
class EclatiraPlayer extends AudioWorkletProcessor {
  constructor(options) {
    super();
    const opts = (options && options.processorOptions) || {};
    this.sourceRate = opts.sourceRate || 24000;
    // If the context did not get 24000, step through the source at a ratio
    // rather than playing it at the wrong speed.
    this.ratio = this.sourceRate / sampleRate;
    this.queue = [];
    this.chunk = null;
    this.position = 0;
    this.draining = false;

    this.port.onmessage = (event) => {
      const data = event.data;
      if (data.type === "audio") {
        this.queue.push(new Int16Array(data.buffer));
        if (!this.draining) {
          this.draining = true;
          this.port.postMessage({ type: "playback_start" });
        }
      } else if (data.type === "flush") {
        // Barge-in. Everything queued belongs to a reply the user has already
        // talked over, so it must be dropped rather than played out.
        this.queue = [];
        this.chunk = null;
        this.position = 0;
        this.draining = false;
      }
    };
  }

  nextSample() {
    while (!this.chunk || this.position >= this.chunk.length) {
      if (!this.queue.length) return null;
      this.chunk = this.queue.shift();
      this.position = 0;
    }
    return this.chunk[Math.floor(this.position)];
  }

  process(inputs, outputs) {
    const channel = outputs[0][0];
    for (let i = 0; i < channel.length; i++) {
      const sample = this.nextSample();
      if (sample === null) {
        channel[i] = 0;
        if (this.draining) {
          this.draining = false;
          this.port.postMessage({ type: "playback_end" });
        }
        continue;
      }
      channel[i] = sample / 32768;
      this.position += this.ratio;
    }
    return true;
  }
}
registerProcessor("eclatira-player", EclatiraPlayer);
```

Wire it up and decode incoming audio:

```js theme={"system"}
/** Start playback and return a handle with push() and flush(). */
async function startPlayback({ onSpeakingChange } = {}) {
  const context = new AudioContext({ sampleRate: 24000 });
  if (context.state === "suspended") await context.resume();

  await context.audioWorklet.addModule("/eclatira-player.js");
  const node = new AudioWorkletNode(context, "eclatira-player", {
    processorOptions: { sourceRate: 24000 },
  });
  node.connect(context.destination);

  node.port.onmessage = ({ data }) => {
    if (data.type === "playback_start") onSpeakingChange?.(true);
    else if (data.type === "playback_end") onSpeakingChange?.(false);
  };

  return {
    context,
    node,
    push(base64) {
      node.port.postMessage({ type: "audio", buffer: fromBase64(base64) });
    },
    flush() {
      node.port.postMessage({ type: "flush" });
    },
  };
}

function fromBase64(value) {
  const binary = atob(value);
  const bytes = new Uint8Array(binary.length);
  for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i);
  return bytes.buffer;
}

/** Route one server message into playback. */
function handleMessage(raw, player) {
  const message = JSON.parse(raw);
  if (message.type) return; // typed control message; see the protocol reference.
  if (message.interrupted) player.flush();
  for (const part of message.parts || []) {
    if (part.type === "audio/pcm") player.push(part.data);
  }
}
```

Dividing by `32768` on playback is correct. Dividing by 32767 would let `-32768` go slightly past −1 and clip.

## Barge-in

When the user talks over the agent, the agent's audio must stop at once. If it does not, both voices overlap and the conversation falls apart.

### Why the server signal is not enough

The server sends `interrupted: true` when it detects the user cutting in. Handle it, but do not rely on it alone:

| Trigger | Speed |
| - | - |
| `interrupted: true` from the server | Slow. It waits for the agent to detect speech and for the message to cross the network. |
| Local speech detection in the recorder | Fast. It runs in the audio thread, before the audio is even sent. |

Your playback queue already holds all the audio the server sent. If you do not clear it, the agent keeps talking until the queue is empty. Muting or ignoring new chunks does not help.

### Detection settings

These values are in the recorder worklet above:

| Setting | Value | Meaning |
| - | - | - |
| `energyThreshold` | `0.015` | Energy level above which a block counts as speech |
| `speechSeconds` | `0.6` | How long speech must last before `speech_start` fires. Filters out clicks and coughs. |
| `silenceSeconds` | `0.8` | How long silence must last before `speech_end` fires. Covers normal pauses. |

Lower values make the agent stop on room noise. Higher values make interruptions feel slow. Test changes with a real microphone in a real room, one at a time.

### Minimum delay

With `speechSeconds` at `0.6`, the queue cannot clear sooner than about **600 ms** after the user starts talking. Measured, it clears at **632 ms**. So:

* About 600 ms of agent speech plays over the user even when everything works. This is expected.
* A very short sound never clears the queue locally. Only the slower server signal catches it.

The local detector only clears playback. It does not decide turns.

### Wire up both triggers

```js theme={"system"}
/**
 * Connect capture and playback so that either trigger flushes the queue.
 * `ws` is an open realtime WebSocket, `stream` a MediaStream with audio.
 */
async function startVoice(ws, stream) {
  const player = await startPlayback({
    onSpeakingChange: (speaking) => console.log("agent speaking:", speaking),
  });

  await startCapture(ws, stream, {
    // Trigger 1: the user started talking. Fires in the audio thread.
    onSpeechStart: () => player.flush(),
    onSpeechEnd: () => {},
  });

  ws.onmessage = (event) => {
    const message = JSON.parse(event.data);
    if (message.type === "error") {
      console.error(message.code, message.error);
      return;
    }
    if (message.type) return;
    // Trigger 2: the server confirmed an interruption. Slower, but catches
    // interruptions the local detector's thresholds missed.
    if (message.interrupted) player.flush();
    for (const part of message.parts || []) {
      if (part.type === "audio/pcm") player.push(part.data);
    }
  };

  return player;
}
```

Keep sending microphone audio while you clear playback. The agent needs it to understand the interruption.

## Video frames

Camera and screen sessions send images the same way, on the same socket as audio. The only difference is what the agent is told it is looking at.

### Format

| Property | Value |
| - | - |
| `mime_type` | `image/jpeg` |
| `data` | Base64 of the JPEG bytes, without the `data:` prefix |
| Image type | One complete JPEG per message. Not a video stream. |
| Other formats | PNG, WebP, AVIF and H.264 are rejected with `unsupported_mime_type`. |
| Order | Any. Mix images and audio freely. |
| Size, quality, color space | Not checked by the server |

`canvas.toDataURL("image/jpeg", q)` returns `data:image/jpeg;base64,...`. Send only the part after the comma. The server removes the prefix if you leave it, but a value with no comma fails with `invalid_base64`.

### Server pacing

| Behavior | Value |
| - | - |
| Forwarding rate | At most one image every **3 seconds** |
| Extra images | The newest waiting image is sent. The others are dropped. |
| Check interval | Every 0.2 seconds, so an image can wait up to 0.2 s after the window opens |
| Cost of sending too many | Bandwidth and CPU only |

The agent always gets the most recent image. Sending one about every 2 seconds means a fresh image is waiting when each window opens. Sending at video speed gains nothing.

### Recommended settings

The server sets no limits, so these are choices, not rules:

| Setting | Library default | Why |
| - | - | - |
| Longest side | `1280` px | Keeps code and UI text readable without huge images |
| JPEG quality | `0.8` | Readable small text at a reasonable size |
| Interval | `2000` ms | The server forwards up to one image per second. Use `1000` for the freshest view. |

<Note>
  Google's own reference client is more careful: 0.5 images per second, scaled to 25%. Move toward those numbers for very large screens or long sessions. Eclatira's defaults favor readable text.
</Note>

### Capture an image

```js theme={"system"}
/**
 * Draw the current video frame to a canvas, JPEG-encode it, and send it.
 * `video` is a playing HTMLVideoElement whose srcObject is the camera or
 * screen MediaStream.
 */
function sendFrame(ws, video, canvas, { maxEdge = 1280, quality = 0.8 } = {}) {
  if (!video.videoWidth || video.readyState < 2) return;

  const scale = Math.min(1, maxEdge / Math.max(video.videoWidth, video.videoHeight));
  canvas.width = Math.round(video.videoWidth * scale);
  canvas.height = Math.round(video.videoHeight * scale);
  canvas.getContext("2d").drawImage(video, 0, 0, canvas.width, canvas.height);

  // toDataURL returns "data:image/jpeg;base64,...." — only the payload after
  // the comma is valid base64.
  const dataUrl = canvas.toDataURL("image/jpeg", quality);
  ws.send(JSON.stringify({
    mime_type: "image/jpeg",
    data: dataUrl.slice(dataUrl.indexOf(",") + 1),
  }));
}

/** Start frame capture. Returns a stop function. */
function startFrames(ws, video, options = {}) {
  const canvas = document.createElement("canvas");
  const tick = () => sendFrame(ws, video, canvas, options);
  tick();
  const timer = setInterval(tick, options.intervalMs ?? 2000);
  return () => clearInterval(timer);
}
```

Keep the `video.readyState < 2` check. Drawing a video before it has a frame gives a black image. The server accepts it, and the agent describes a black screen.

### Video with no audio

<Warning>
  **A camera or screen session with no microphone audio never gets a reply.** The session connects, accepts every image, and nothing else happens.
</Warning>

The server sends this once, when an image arrives at least 8 seconds after the first image and no audio has arrived:

```json theme={"system"}
{
  "type": "error",
  "code": "no_audio_received",
  "error": "Frames are arriving but no audio is. The agent replies to speech, so a camera or screen session must also stream 'audio/pcm' from the microphone (or send a 'text/plain' message) before it will respond."
}
```

The session stays open. This check only runs when an image arrives. If you stop sending images, you never get it. So no error does not prove audio is arriving.

On screen sessions, the usual cause is `getDisplayMedia`:

```js theme={"system"}
// WRONG. getDisplayMedia's audio track is system or tab audio, never the
// microphone. The agent hears the shared tab and never hears the user.
const stream = await navigator.mediaDevices.getDisplayMedia({ video: true, audio: true });
```

```js theme={"system"}
// CORRECT. Acquire the two separately and merge the tracks.
const display = await navigator.mediaDevices.getDisplayMedia({
  video: { width: { ideal: 1280 }, height: { ideal: 720 } },
  audio: false,
});
const mic = await navigator.mediaDevices.getUserMedia({ audio: true, video: false });
const stream = new MediaStream([...display.getVideoTracks(), ...mic.getAudioTracks()]);

// The browser's own "Stop sharing" control ends the video track without
// touching the WebSocket. Nothing else tells you the share is over.
display.getVideoTracks()[0].addEventListener("ended", () => {
  /* end the session and stop every track */
});
```

### Message size limit

The only size limit is the server's WebSocket limit: **16 MiB** per message, including the base64 overhead. That means about 12 MiB of JPEG.

A message over the limit closes the connection with **no error message**. A 1280 px JPEG at quality 0.8 is usually 100 to 400 KB, far below the limit.

## The message envelope

Every message you send is a JSON object in a WebSocket **text** frame, with two fields:

```json theme={"system"}
{ "mime_type": "audio/pcm", "data": "<base64>" }
```

| `mime_type` | `data` |
| - | - |
| `audio/pcm` | Base64 of raw PCM: 16 kHz, 16-bit signed little-endian, mono, no header |
| `image/jpeg` | Base64 of the JPEG bytes |
| `text/plain` | **Plain UTF-8, not base64** |

### Other accepted spellings

Write `audio/pcm`, `image/jpeg` and `text/plain`. The server also accepts these, so a near-miss does not fail silently:

| Accepted | Examples |
| - | - |
| Parameters (removed before matching) | `audio/pcm;rate=16000`, `audio/pcm; codecs=1` |
| Other audio names | `audio/l16`, `audio/x-raw`, `audio/raw` |
| Other image name | `image/jpg` |
| Other text name | `text` |
| Any capitalization | `Audio/PCM`, `IMAGE/JPEG` |
| Spaces around the value | `" audio/pcm "` |
| A `data:` prefix on `data` | `data:image/jpeg;base64,AAAA...` |
| Base64 without `=` padding | `AAA` for `AAA=` |

### The rate parameter is ignored

<Warning>
  `audio/pcm;rate=48000` does **not** make the server resample. The server removes the rate and always treats the audio as 16 kHz. Resample in your client.
</Warning>

### Binary frames

Binary WebSocket frames are **not supported**. Each one gets this reply:

```json theme={"system"}
{
  "type": "error",
  "code": "binary_frame_unsupported",
  "error": "Binary WebSocket frames are not supported. Send a JSON text frame with 'mime_type' and base64 'data'."
}
```

The session stays open, but the frame is dropped. A client that only sends binary audio never gets a reply.

For every error code, see [Errors and close codes](/realtime/errors#error-messages).

## Silent failures

Every row below leaves the WebSocket open and healthy. Unless the row says so, there is **no error**.

| Mistake | What you see | Fix |
| - | - | - |
| Audio at 48000 or 44100 Hz | Transcript empty or nonsense. Agent never speaks. | Check `context.sampleRate` and [resample](#the-resampling-worklet). |
| Audio at 24000 Hz | Same. Sometimes a garbled transcript. | 16000 in, 24000 out. |
| `audio/pcm;rate=48000` instead of resampling | Same. | [The rate is ignored.](#the-rate-parameter-is-ignored) Resample. |
| Float32 samples sent as-is | Agent never speaks. The server hears loud noise. | Convert with [`floatToInt16`](#convert-float32-to-int16). |
| 8-bit or 32-bit samples | Agent never speaks, or replies to noise. | 16-bit signed only. |
| Unsigned 16-bit samples | Agent never speaks. | Signed, −32768 to 32767. |
| Big-endian samples | Agent never speaks, or hears nonsense. | Little-endian. In Python use `"<h"`, not `">h"`. |
| Stereo | Agent never speaks. | Mono. Use `channelCount: 1` or take channel 0. |
| WAV header on every chunk | Clicks. Agent never speaks. | Raw samples only. |
| WAV header on the first chunk only | Usually works, with one click. | Remove it anyway. |
| `data:` prefix left on an image | Works. | Remove it anyway. A value with no comma fails with `invalid_base64`. |
| Binary frames | One `binary_frame_unsupported` error per frame. | JSON text frames only. |
| Camera or screen with no microphone | Nothing happens. One `no_audio_received` after about 8 s. | Send `audio/pcm`, or send `text/plain`. |
| `getDisplayMedia({ audio: true })` used as the microphone | Agent hears the computer's sound, not the user. | [Get two streams and merge them.](#video-with-no-audio) |
| Audio only sent while the user speaks | Agent interrupts or never answers. | Send all the time, including silence. |
| Images sent at video speed | Works, but wastes bandwidth. | Send every 1.5 to 3 seconds. |
| Playback not cleared on barge-in | Agent talks over the user. | [Clear on both triggers.](#wire-up-both-triggers) |
| One `AudioContext` for both directions | Agent sounds slow, or never replies. | [Use two.](#you-need-two-audiocontexts) |
| Image drawn before the video loaded | Agent describes a black screen. | Check `video.videoWidth` and `video.readyState`. |
| User clicks the browser's **Stop sharing** | Images stop. Session keeps running and billing. | Listen for the video track's `ended` event. |
| Message over 16 MiB | Connection closes. No error. | [Scale images down.](#recommended-settings) |

## Outside the browser

With `ffmpeg`, the format maps to these flags:

```bash theme={"system"}
# Any input file to exactly what the realtime socket accepts:
# s16le = signed 16-bit little-endian, ar = 16000 Hz, ac = 1 channel,
# and the `s16le` muxer writes raw samples with no header.
ffmpeg -i input.wav -f s16le -ar 16000 -ac 1 output.raw
```

In Python, always pack little-endian explicitly:

```python theme={"system"}
import base64
import struct

def encode_pcm16(samples: list[float]) -> str:
    """Encode float samples in [-1, 1] as base64 16-bit little-endian PCM."""
    frames = bytearray()
    for sample in samples:
        clamped = max(-1.0, min(1.0, sample))
        value = int(clamped * 0x8000 if clamped < 0 else clamped * 0x7FFF)
        frames += struct.pack("<h", value)  # "<" is little-endian, "h" is int16.
    return base64.b64encode(bytes(frames)).decode("ascii")
```

## Next steps

<CardGroup cols={2}>
  <Card title="WebSocket protocol" icon="plug" href="/realtime-protocol">
    Every message the server sends.
  </Card>

  <Card title="Troubleshooting" icon="bug" href="/realtime/troubleshooting">
    Step-by-step checks for a silent agent.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.