> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# WebSocket protocol

> The messages a web session sends and receives, how a connection starts, and how turns work.

This page describes the WebSocket protocol behind web sessions. It covers sessions created with an API key through `POST /api/v1/realtime/sessions`.

<Info>
  Building in a browser? The [client library](/realtime/browser-voice) implements all of this for you. For the exact audio and image formats, see [Media format](/realtime/media-format).
</Info>

## Connect

<Steps>
  <Step title="Create a session">
    Call `POST /api/v1/realtime/sessions` from your server. See the [API reference](/api-reference/realtime/create-a-realtime-session).
  </Step>

  <Step title="Open ws_url within 60 seconds">
    Open it exactly as returned. The token works for one connection attempt only, even if that attempt fails.
  </Step>

  <Step title="Wait for session_created">
    The server sends [`session_created`](#session_created) first. It may then send a one-time [`credits_low`](#credits_low). After that, the conversation starts.
  </Step>
</Steps>

The URL looks like this:

```
wss://app.eclatira.com/ws/{user_id}?session_token=...&source=...
```

| Part | Where | Required | Description |
| - | - | - | - |
| `user_id` | Path | Yes | Must match the `user_id` the session was created with. Use a new value for each session. |
| `session_token` | Query | Yes | The single-use token from the create call. |
| `source` | Query | No | Shown for information only. The session type was fixed when the session was created. |

<Warning>
  **Do not use your app's user ID as `user_id`.** Two sessions open at the same time with the same `user_id` overwrite each other on the server.

  For `text` sessions, reusing a recent `user_id` has another effect. If the old session is still in memory (up to 30 minutes), the new connection continues that conversation and `session_created` reports `resumed: true`.

  Generate a new value for each session. A UUID works. So does `{your_user_id}:{uuid}` if you want your own ID in the logs.
</Warning>

## Try it

The example below uses a `text` session. It needs no audio code, so it tests only the connection, the token and the message format. Text messages are plain UTF-8, not base64. The agent replies right away.

Create a session with `"source": "text"`, then run:

```python theme={"system"}
# pip install websockets
import asyncio
import json

import websockets

WS_URL = "wss://app.eclatira.com/ws/YOUR_USER_ID?session_token=YOUR_TOKEN&source=text"


async def main() -> None:
    async with websockets.connect(WS_URL) as ws:
        created = json.loads(await ws.recv())
        print(created)  # {"type": "session_created", "session_id": "...", "resumed": false}

        await ws.send(json.dumps({"mime_type": "text/plain", "data": "Hello!"}))

        while True:
            message = json.loads(await ws.recv())
            for part in message.get("parts", []):
                if part.get("type") == "text":
                    print(part["data"], end="", flush=True)
            if message.get("turn_complete"):
                print()
                break


asyncio.run(main())
```

If the agent has a first message, the first turn you read is that greeting. Keep reading to see the reply.

## Turn-taking

There are no turn messages. Nothing you send starts or ends a turn.

* **Audio drives turns.** The agent listens to the `audio/pcm` stream and decides when the user has finished. Send audio all the time, including silence. Messages like `commit` or `end_of_turn` do not exist and are rejected.
* **Text starts a turn.** A `text/plain` message on any session makes the agent reply right away. It is the only way to get a reply without speech.

<Warning>
  **Camera and screen sessions need audio too.** Images do not start a turn. A session that sends only images connects, accepts them all, and never replies.

  The server sends one [`no_audio_received`](/realtime/errors#error-messages) error when an image arrives at least 8 seconds after the first image and no audio has arrived. It only checks when an image arrives, so if you stop sending, you never get it.
</Warning>

When the user talks over the agent, the server sends a [voice turn event](#voice-turn-event) with `interrupted: true`. This arrives late. By then, buffered agent audio is already playing. Detect speech locally and clear your playback queue. See [Barge-in](/realtime/media-format#barge-in).

## Image pacing

Images use the same socket as audio. Mix them in any order. The server paces them:

* It forwards **at most one image per second**.
* If you send more, it keeps the **newest** waiting image and drops the rest.
* It checks every 200 ms, so an image can wait up to about 200 ms after the window opens.

<Tip>
  Send one image about every second. Then the agent's view is never more than about a second old.
</Tip>

The server sets no limit on image size, quality or color space. The only limit is 16 MiB per WebSocket message. A longest side of 1280 px at quality 0.8 keeps screen text readable.

## Voice and text sessions differ

Voice sessions (`audio`, `camera` and `screen`) and `text` sessions receive different messages:

| Message | Voice sessions | Text sessions |
| - | - | - |
| [`session_created`](#session_created) | Yes | Yes |
| [`session_ended`](#session_ended) | Yes | Yes |
| [`credits_low`](#credits_low) | Yes | No |
| [`error`](#error) | All codes | Only `binary_frame_unsupported` and `missing_data` |
| [Voice turn event](#voice-turn-event) | Yes | No |
| [Voice content event](#voice-content-event) | Yes | No |
| [Text turn event](#text-turn-event) | No | Yes |
| [`topup_succeeded`](#topup_succeeded) | No | Yes |
| [`limit_reached`](#limit_reached) | No | Yes |

This is because credits work differently. Voice sessions check credits once, before they start, and are never cut off. Text sessions check after every turn and can stop mid-chat.

## Messages you send

Every message is a JSON object in a WebSocket **text** frame. There is only one shape:

| Field | Type | Required | Description |
| - | - | - | - |
| `mime_type` | `"audio/pcm"`, `"image/jpeg"` or `"text/plain"` | Yes | Camera and screen both use `image/jpeg`. |
| `data` | string | Yes | `audio/pcm`: base64 of raw 16 kHz, 16-bit signed little-endian mono PCM, no header. `image/jpeg`: base64 of the JPEG. `text/plain`: plain UTF-8, **not** base64. |

```json theme={"system"}
{ "mime_type": "text/plain", "data": "Hello!" }
```

The agent's audio comes back at a different rate: 24 kHz. You need one `AudioContext` for recording and another for playback.

<Warning>
  **Audio at the wrong sample rate fails silently.** Browsers often ignore `new AudioContext({ sampleRate: 16000 })`. The server cannot tell. Check `context.sampleRate` and resample. See [Media format](/realtime/media-format#the-audiocontext-sample-rate-problem).
</Warning>

### Accepted spellings

Write `audio/pcm`, `image/jpeg` and `text/plain`. The server also accepts:

| You send | Treated as |
| - | - |
| `audio/pcm;rate=16000` | `audio/pcm`. This is what the google-genai SDK sends. The rate is ignored. |
| `audio/l16`, `audio/x-raw`, `audio/raw` | `audio/pcm` |
| `image/jpg` | `image/jpeg` |
| `text` | `text/plain` |
| Any capitalization, extra spaces | The normal spelling |

For `data`, the server also removes a `data:...;base64,` prefix and adds missing `=` padding.

### Binary frames

Binary frames are not supported on any session. The server replies with a `binary_frame_unsupported` error and keeps the session open.

### Text sessions ignore media

A `text` session only reads `text/plain`. It ignores `audio/pcm` and `image/jpeg` with no error. Invalid JSON does not produce an error either. Instead, the agent sends a normal-looking apology message. Check your own JSON on text sessions.

## Messages you receive

<Note>
  **Check `type` first.** Session, billing and error messages have a top-level `type`. Voice turn and content events do not. Identify those by their fields: `turn_complete`, `is_partial`, `parts`, `input_transcription`, `output_transcription`.

  Entries inside `parts` also have a `type`. That one names the part, not the message.
</Note>

### session\_created

Sent once, right after the connection succeeds. All session types.

| Field | Type | Description |
| - | - | - |
| `type` | `"session_created"` | |
| `session_id` | string | The session's ID. Use it in your logs and support requests. |
| `user_id` | string | The `user_id` from the URL. |
| `resumed` | boolean | `true` if this continued an earlier text session. Always `false` on voice sessions. |

```json theme={"system"}
{ "type": "session_created", "session_id": "abc123", "user_id": "session-7f3c1e9a", "resumed": false }
```

### session\_ended

Sent when the agent ends the conversation. All session types. Voice sessions send it after a 5 second delay. Text sessions send it 3 seconds after the goodbye message.

<Warning>
  **On voice sessions, the server does not close the socket after this.** Call `ws.close()` yourself when you get `session_ended`. Otherwise the socket stays open and holds one of your session slots.

  On text sessions, the server closes the socket for you, with code `1000`.
</Warning>

| Field | Type | Description |
| - | - | - |
| `type` | `"session_ended"` | |
| `reason` | string | For example `"agent_ended"`, or a reason the agent gave. |

```json theme={"system"}
{ "type": "session_ended", "reason": "agent_ended" }
```

### credits\_low

Sent once, right after `session_created`, if the workspace had already used 90% or more of its plan. Voice sessions only.

| Field | Type | Description |
| - | - | - |
| `type` | `"credits_low"` | |
| `pct_used` | integer | Percent of the plan already used, for example `92`. |

```json theme={"system"}
{ "type": "credits_low", "pct_used": 92 }
```

### error

Sent when the server cannot use something you sent. **The session stays open.** Keep reading and keep sending.

| Field | Type | Description |
| - | - | - |
| `type` | `"error"` | |
| `code` | string | A stable code. Match on this. |
| `error` | string | A readable explanation. The wording can change. |

```json theme={"system"}
{
  "type": "error",
  "code": "no_audio_received",
  "error": "Frames are arriving but no audio is. The agent replies to speech, so a camera or screen session must also stream 'audio/pcm' from the microphone (or send a 'text/plain' message) before it will respond."
}
```

For every code, its cause and its fix, see [Errors and close codes](/realtime/errors#error-messages).

### Voice turn event

Voice sessions only. Marks the end of an agent turn, or an interruption. It has no content. `parts` is always empty and both transcriptions are always `null`.

| Field | Type | Description |
| - | - | - |
| `turn_complete` | boolean | |
| `interrupted` | boolean | `true` if the user talked over the agent. Clear queued agent audio. |
| `parts` | array | Always empty. |
| `input_transcription` | `null` | |
| `output_transcription` | `null` | |

```json theme={"system"}
{ "turn_complete": true, "interrupted": false, "parts": [], "input_transcription": null, "output_transcription": null }
```

### Voice content event

Voice sessions only. Carries agent audio, tool calls and live transcripts. At least one of `parts`, `input_transcription` or `output_transcription` has content.

| Field | Type | Description |
| - | - | - |
| `is_partial` | boolean | `false` on the last chunk of a turn. |
| `parts` | array of [voice parts](#voice-parts) | Often empty on transcript-only updates. |
| `input_transcription` | [Transcription](#transcription) or `null` | What the user said. |
| `output_transcription` | [Transcription](#transcription) or `null` | What the agent said. |

```json theme={"system"}
{
  "is_partial": false,
  "parts": [{ "type": "audio/pcm", "data": "<base64>" }],
  "input_transcription": { "text": "What's the weather like?", "is_final": true },
  "output_transcription": null
}
```

The agent's words arrive in `output_transcription`, never as a text part.

#### Voice parts

Each part has a `type`:

| `type` | Fields | Description |
| - | - | - |
| `audio/pcm` | `data` | Base64 agent audio: 24 kHz, 16-bit signed little-endian mono PCM. Chunks are continuous. Queue them into one stream. |
| `function_call` | `data.name`, `data.args` | The agent called a tool. |
| `function_response` | `data.name`, `data.response` | The tool returned. |

#### Transcription

| Field | Type |
| - | - |
| `text` | string |
| `is_final` | boolean |

Transcripts send the full text so far, not just new words. Replace the current line with each update. The line is complete when `is_final` is `true`.

### Text turn event

Text sessions only. Carries the agent's streamed reply, its tool calls, and the first greeting. Tool calls come in a separate message after the text.

| Field | Type | Description |
| - | - | - |
| `is_partial` | boolean | |
| `parts` | array of text parts | |
| `turn_complete` | boolean | |

Text parts have a `type`:

| `type` | Fields |
| - | - |
| `text` | `data` |
| `function_call` | `data.name`, `data.args` |
| `function_response` | `data.name`, `data.response` |

```json theme={"system"}
{ "is_partial": true, "parts": [{ "type": "text", "data": "Sure, I can help with that." }], "turn_complete": false }
```

### topup\_succeeded

Text sessions only. Sent when a turn reaches the plan limit and credits are topped up automatically.

| Field | Type |
| - | - |
| `type` | `"topup_succeeded"` |
| `overage_credits_added` | integer |

```json theme={"system"}
{ "type": "topup_succeeded", "overage_credits_added": 5000 }
```

### limit\_reached

Text sessions only. Sent when the plan limit is reached and no top-up was possible. The connection stays open, but the agent stops replying.

Voice sessions never get this. A voice session over its limit still finishes. The extra usage moves to the next billing cycle.

The message has a `type` and also the same fields as a text turn event, so it shows as a normal chat message.

| Field | Type | Description |
| - | - | - |
| `type` | `"limit_reached"` | |
| `topup_attempted` | boolean | `true` if a top-up was tried and failed. |
| `is_partial` | boolean | Always `false`. |
| `parts` | array of text parts | One text part with an explanation. |
| `turn_complete` | boolean | Always `true`. |

```json theme={"system"}
{
  "type": "limit_reached",
  "topup_attempted": true,
  "is_partial": false,
  "parts": [{ "type": "text", "data": "You've reached your plan's credit limit for this month." }],
  "turn_complete": true
}
```

## Close codes

A refused connection is accepted and then closed with a code. The per-IP rate limit is the exception. It refuses the handshake itself, so some clients see a failed handshake instead of `4029`.

| Code | Meaning |
| - | - |
| `4001` | Invalid, expired or already-used token |
| `4029` | Too many connections from one IP, or too many sessions for your plan |
| `4402` | Subscription past due or unpaid |
| `4028` | Plan limit reached, or not enough credit for a voice session |
| `4003` | Camera or screen share was turned off for the agent |
| `1011` | Server error starting the session |
| `1000` | Normal close |

See [Errors and close codes](/realtime/errors#close-codes) for causes and fixes.

## Save the transcript

You cannot fetch a web session later through `GET /api/v1/calls/{call_id}`. That endpoint only finds phone calls. `session_id` is useful for your logs, but no endpoint accepts it.

To keep a record:

* **Collect it during the session.** On voice sessions, append `input_transcription.text` and `output_transcription.text` from each voice content event. Treat `is_final: true` as the settled text. On text sessions, append the `text` parts of each text turn event. Send it to your backend as it grows.
* **Use a webhook.** A registered [webhook](/webhooks) gets `call.completed` with the full transcript when the session ends.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.