Skip to main content
This page describes the WebSocket protocol behind web sessions. It covers sessions created with an API key through POST /api/v1/realtime/sessions.
Building in a browser? The client library implements all of this for you. For the exact audio and image formats, see Media format.

Connect

1

Create a session

Call POST /api/v1/realtime/sessions from your server. See the API reference.
2

Open ws_url within 60 seconds

Open it exactly as returned. The token works for one connection attempt only, even if that attempt fails.
3

Wait for session_created

The server sends session_created first. It may then send a one-time credits_low. After that, the conversation starts.
The URL looks like this:
Do not use your app’s user ID as user_id. Two sessions open at the same time with the same user_id overwrite each other on the server.For text sessions, reusing a recent user_id has another effect. If the old session is still in memory (up to 30 minutes), the new connection continues that conversation and session_created reports resumed: true.Generate a new value for each session. A UUID works. So does {your_user_id}:{uuid} if you want your own ID in the logs.

Try it

The example below uses a text session. It needs no audio code, so it tests only the connection, the token and the message format. Text messages are plain UTF-8, not base64. The agent replies right away. Create a session with "source": "text", then run:
If the agent has a first message, the first turn you read is that greeting. Keep reading to see the reply.

Turn-taking

There are no turn messages. Nothing you send starts or ends a turn.
  • Audio drives turns. The agent listens to the audio/pcm stream and decides when the user has finished. Send audio all the time, including silence. Messages like commit or end_of_turn do not exist and are rejected.
  • Text starts a turn. A text/plain message on any session makes the agent reply right away. It is the only way to get a reply without speech.
Camera and screen sessions need audio too. Images do not start a turn. A session that sends only images connects, accepts them all, and never replies.The server sends one no_audio_received error when an image arrives at least 8 seconds after the first image and no audio has arrived. It only checks when an image arrives, so if you stop sending, you never get it.
When the user talks over the agent, the server sends a voice turn event with interrupted: true. This arrives late. By then, buffered agent audio is already playing. Detect speech locally and clear your playback queue. See Barge-in.

Image pacing

Images use the same socket as audio. Mix them in any order. The server paces them:
  • It forwards at most one image per second.
  • If you send more, it keeps the newest waiting image and drops the rest.
  • It checks every 200 ms, so an image can wait up to about 200 ms after the window opens.
Send one image about every second. Then the agent’s view is never more than about a second old.
The server sets no limit on image size, quality or color space. The only limit is 16 MiB per WebSocket message. A longest side of 1280 px at quality 0.8 keeps screen text readable.

Voice and text sessions differ

Voice sessions (audio, camera and screen) and text sessions receive different messages: This is because credits work differently. Voice sessions check credits once, before they start, and are never cut off. Text sessions check after every turn and can stop mid-chat.

Messages you send

Every message is a JSON object in a WebSocket text frame. There is only one shape:
The agent’s audio comes back at a different rate: 24 kHz. You need one AudioContext for recording and another for playback.
Audio at the wrong sample rate fails silently. Browsers often ignore new AudioContext({ sampleRate: 16000 }). The server cannot tell. Check context.sampleRate and resample. See Media format.

Accepted spellings

Write audio/pcm, image/jpeg and text/plain. The server also accepts: For data, the server also removes a data:...;base64, prefix and adds missing = padding.

Binary frames

Binary frames are not supported on any session. The server replies with a binary_frame_unsupported error and keeps the session open.

Text sessions ignore media

A text session only reads text/plain. It ignores audio/pcm and image/jpeg with no error. Invalid JSON does not produce an error either. Instead, the agent sends a normal-looking apology message. Check your own JSON on text sessions.

Messages you receive

Check type first. Session, billing and error messages have a top-level type. Voice turn and content events do not. Identify those by their fields: turn_complete, is_partial, parts, input_transcription, output_transcription.Entries inside parts also have a type. That one names the part, not the message.

session_created

Sent once, right after the connection succeeds. All session types.

session_ended

Sent when the agent ends the conversation. All session types. Voice sessions send it after a 5 second delay. Text sessions send it 3 seconds after the goodbye message.
On voice sessions, the server does not close the socket after this. Call ws.close() yourself when you get session_ended. Otherwise the socket stays open and holds one of your session slots.On text sessions, the server closes the socket for you, with code 1000.

credits_low

Sent once, right after session_created, if the workspace had already used 90% or more of its plan. Voice sessions only.

error

Sent when the server cannot use something you sent. The session stays open. Keep reading and keep sending.
For every code, its cause and its fix, see Errors and close codes.

Voice turn event

Voice sessions only. Marks the end of an agent turn, or an interruption. It has no content. parts is always empty and both transcriptions are always null.

Voice content event

Voice sessions only. Carries agent audio, tool calls and live transcripts. At least one of parts, input_transcription or output_transcription has content.
The agent’s words arrive in output_transcription, never as a text part.

Voice parts

Each part has a type:

Transcription

Transcripts send the full text so far, not just new words. Replace the current line with each update. The line is complete when is_final is true.

Text turn event

Text sessions only. Carries the agent’s streamed reply, its tool calls, and the first greeting. Tool calls come in a separate message after the text. Text parts have a type:

topup_succeeded

Text sessions only. Sent when a turn reaches the plan limit and credits are topped up automatically.

limit_reached

Text sessions only. Sent when the plan limit is reached and no top-up was possible. The connection stays open, but the agent stops replying. Voice sessions never get this. A voice session over its limit still finishes. The extra usage moves to the next billing cycle. The message has a type and also the same fields as a text turn event, so it shows as a normal chat message.

Close codes

A refused connection is accepted and then closed with a code. The per-IP rate limit is the exception. It refuses the handshake itself, so some clients see a failed handshake instead of 4029. See Errors and close codes for causes and fixes.

Save the transcript

You cannot fetch a web session later through GET /api/v1/calls/{call_id}. That endpoint only finds phone calls. session_id is useful for your logs, but no endpoint accepts it. To keep a record:
  • Collect it during the session. On voice sessions, append input_transcription.text and output_transcription.text from each voice content event. Treat is_final: true as the settled text. On text sessions, append the text parts of each text turn event. Send it to your backend as it grows.
  • Use a webhook. A registered webhook gets call.completed with the full transcript when the session ends.