POST /api/v1/realtime/sessions.
Building in a browser? The client library implements all of this for you. For the exact audio and image formats, see Media format.
Connect
1
Create a session
Call
POST /api/v1/realtime/sessions from your server. See the API reference.2
Open ws_url within 60 seconds
Open it exactly as returned. The token works for one connection attempt only, even if that attempt fails.
3
Wait for session_created
The server sends
session_created first. It may then send a one-time credits_low. After that, the conversation starts.Try it
The example below uses atext session. It needs no audio code, so it tests only the connection, the token and the message format. Text messages are plain UTF-8, not base64. The agent replies right away.
Create a session with "source": "text", then run:
Turn-taking
There are no turn messages. Nothing you send starts or ends a turn.- Audio drives turns. The agent listens to the
audio/pcmstream and decides when the user has finished. Send audio all the time, including silence. Messages likecommitorend_of_turndo not exist and are rejected. - Text starts a turn. A
text/plainmessage on any session makes the agent reply right away. It is the only way to get a reply without speech.
interrupted: true. This arrives late. By then, buffered agent audio is already playing. Detect speech locally and clear your playback queue. See Barge-in.
Image pacing
Images use the same socket as audio. Mix them in any order. The server paces them:- It forwards at most one image per second.
- If you send more, it keeps the newest waiting image and drops the rest.
- It checks every 200 ms, so an image can wait up to about 200 ms after the window opens.
Voice and text sessions differ
Voice sessions (audio, camera and screen) and text sessions receive different messages:
This is because credits work differently. Voice sessions check credits once, before they start, and are never cut off. Text sessions check after every turn and can stop mid-chat.
Messages you send
Every message is a JSON object in a WebSocket text frame. There is only one shape:AudioContext for recording and another for playback.
Accepted spellings
Writeaudio/pcm, image/jpeg and text/plain. The server also accepts:
For
data, the server also removes a data:...;base64, prefix and adds missing = padding.
Binary frames
Binary frames are not supported on any session. The server replies with abinary_frame_unsupported error and keeps the session open.
Text sessions ignore media
Atext session only reads text/plain. It ignores audio/pcm and image/jpeg with no error. Invalid JSON does not produce an error either. Instead, the agent sends a normal-looking apology message. Check your own JSON on text sessions.
Messages you receive
Check
type first. Session, billing and error messages have a top-level type. Voice turn and content events do not. Identify those by their fields: turn_complete, is_partial, parts, input_transcription, output_transcription.Entries inside parts also have a type. That one names the part, not the message.session_created
Sent once, right after the connection succeeds. All session types.session_ended
Sent when the agent ends the conversation. All session types. Voice sessions send it after a 5 second delay. Text sessions send it 3 seconds after the goodbye message.credits_low
Sent once, right aftersession_created, if the workspace had already used 90% or more of its plan. Voice sessions only.
error
Sent when the server cannot use something you sent. The session stays open. Keep reading and keep sending.Voice turn event
Voice sessions only. Marks the end of an agent turn, or an interruption. It has no content.parts is always empty and both transcriptions are always null.
Voice content event
Voice sessions only. Carries agent audio, tool calls and live transcripts. At least one ofparts, input_transcription or output_transcription has content.
output_transcription, never as a text part.
Voice parts
Each part has atype:
Transcription
Transcripts send the full text so far, not just new words. Replace the current line with each update. The line is complete when
is_final is true.
Text turn event
Text sessions only. Carries the agent’s streamed reply, its tool calls, and the first greeting. Tool calls come in a separate message after the text.
Text parts have a
type:
topup_succeeded
Text sessions only. Sent when a turn reaches the plan limit and credits are topped up automatically.limit_reached
Text sessions only. Sent when the plan limit is reached and no top-up was possible. The connection stays open, but the agent stops replying. Voice sessions never get this. A voice session over its limit still finishes. The extra usage moves to the next billing cycle. The message has atype and also the same fields as a text turn event, so it shows as a normal chat message.
Close codes
A refused connection is accepted and then closed with a code. The per-IP rate limit is the exception. It refuses the handshake itself, so some clients see a failed handshake instead of4029.
See Errors and close codes for causes and fixes.
Save the transcript
You cannot fetch a web session later throughGET /api/v1/calls/{call_id}. That endpoint only finds phone calls. session_id is useful for your logs, but no endpoint accepts it.
To keep a record:
- Collect it during the session. On voice sessions, append
input_transcription.textandoutput_transcription.textfrom each voice content event. Treatis_final: trueas the settled text. On text sessions, append thetextparts of each text turn event. Send it to your backend as it grows. - Use a webhook. A registered webhook gets
call.completedwith the full transcript when the session ends.