session_created arrives, audio goes out, and nothing comes back. There is no error.
This happens because the server cannot check your audio. Raw audio has no header. So 48 kHz audio looks the same as 16 kHz audio. The server accepts it, and the agent never hears anything it understands.
Add logging first
Attach this helper right afternew WebSocket(...). It logs session_created, every error, the close code, and how many bytes of audio you send per second.
32.0 kB/s PCM while the microphone is on. Any other number points to the bug.
The agent never replies
Go through these checks in order.Did session_created arrive?
session_created first, once, on every session.- It never arrived. The server closed the socket during setup. Read
event.codeinoncloseand look it up in Close codes. Do not retry with the same token. It is already used, so you will get4001and hide the real cause. - It arrived. Go to the next step.
Did you get any error messages?
{"type": "error", "code": "..."} for anything it cannot use. The session stays open.no_audio_received: you are sending images but no audio. Go to step 6.- Any other code: look it up in Error messages.
- No errors: the server accepts your messages. The problem is inside the audio itself. Go to step 3.
text sessions do not send these errors. They ignore audio and images without saying anything.Is audio being sent?
AudioContext created outside a click starts suspended. It sends nothing until resume() finishes, with no error.Send audio all the time, including silence. The agent uses the silence to tell when the user has stopped talking.Is the sample rate right?
new AudioContext({ sampleRate: 16000 }) is only a request. Safari and Firefox often ignore it and use 44.1 or 48 kHz. Nothing tells you.Check what you actually got:Is the format right?
- A WAV header. You are probably encoding whole files instead of streaming.
- Float32 samples. The byte rate shows
64.0 kB/sat 16 kHz. - Big-endian samples. Only happens if you build bytes with
DataView.setInt16and leave outlittleEndian. It defaults to big-endian. - Stereo. Two channels read as one fast, garbled channel.
Camera or screen: are you sending microphone audio?
no_audio_received.For screen share, the audio from getDisplayMedia is the computer’s sound, not the microphone. Get the two separately and merge them:Does a text message get a reply?
text/plain message makes the agent reply right away, with no audio involved. It works on every session type.text/plain data as plain text, not base64. If you base64 it, the agent reads the code out loud.ws_url within 60 seconds:Other problems
The agent talks over me when I interrupt
The agent talks over me when I interrupt
0.015, 0.6 seconds of speech to start, and 0.8 seconds of silence to stop.The server also sends interrupted: true. It arrives later, so use it as a backup. See Barge-in.With these settings, the agent keeps talking for about 600 ms after the user starts. A very short word, like “wait”, may not interrupt it. You can lower speechSeconds to react faster. Background noise will then interrupt the agent more often.The agent's audio is choppy or clicks
The agent's audio is choppy or clicks
- One
AudioBufferSourceNodeper chunk adds a gap between chunks. Use one continuous queue instead. - Clicks between chunks mean chunks are lost or out of order. Play them in the order they arrive.
- Clicks all the time mean clipping. Clamp Float32 samples to
[-1, 1]before you convert.
AudioContext at 24000 Hz.The agent hears itself and keeps interrupting itself
The agent hears itself and keeps interrupting itself
- Turn on echo cancellation:
getUserMedia({ audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true } }). - Test with headphones. If the problem goes away, it is echo.
- Check
input_transcription. If it contains the agent’s own words, this is the cause.
The agent sounds too fast, too slow, or garbled
The agent sounds too fast, too slow, or garbled
- Agent too fast and high: you play 24 kHz audio in a faster context without adjusting. Create the playback context at 24000.
- Agent too slow and deep: you play 24 kHz audio as if it were 16 kHz.
- Agent mishears you: your input audio is at the wrong rate or distorted. See steps 4 and 5.
- Both wrong: you use one
AudioContextfor recording and playback. Use two.
It works on localhost but breaks when deployed
It works on localhost but breaks when deployed
http://localhost:3000, but not from your domain. So creating sessions from the browser works locally and fails with a CORS error when deployed. Create sessions on your server. See Authentication.The page is not HTTPS. The microphone and screen capture only work on secure pages. On a deployed http:// page, navigator.mediaDevices is undefined. Serve over HTTPS.429 when creating sessions while debugging
429 when creating sessions while debugging
- Creating a session when a component mounts. React Strict Mode doubles this in development.
- Creating a session on every hot reload.
- Reconnecting in a loop after every close.
start() rejects with invalid_token
start() rejects with invalid_token
4001. Usually the token was already used, often by a double-run effect, a hot reload, or a retry. It may also have expired after 60 seconds. See 4001: invalid token.4029 every time I reload the page
4029 every time I reload the page
- “Too many connections”, or
1006with no reason: too many attempts from your IP. Wait a minute. - “Too many concurrent sessions for your plan”: old sessions still hold slots. Always call
stop()onbeforeunload. A session closed without cleanup can hold its slot for up to 2 hours.
user_id for each session, and reuse it only for retries of that same session.The permission prompt never appears
The permission prompt never appears
- Not HTTPS.
navigator.mediaDevicesdoes not exist on insecure pages. - No user action.
getDisplayMediamust be called directly in a click handler. - Blocked before. If the user blocked the microphone once, the browser remembers. They must reset it in site settings.
- Inside an iframe. The iframe needs
allow="microphone; camera; display-capture". - Permissions-Policy. A strict
Permissions-Policyheader on your page blocks it. - No device. With no microphone, you get
NotFoundError.
name. NotAllowedError, NotFoundError, InvalidStateError and TypeError each point to a different cause.Known limits
These are not bugs in your code.- You cannot fetch a web session through
GET /api/v1/calls/{call_id}. Collect the transcript in the browser, or use a webhook. - A dropped session cannot be resumed. Create a new session and connect again.
- No keepalive is needed. There is no idle timeout or maximum length to handle. If sessions end after the same time every time, check for a proxy or load balancer timeout on your side.
Still stuck?
Collect these before you ask for help:Close code and reason
The first error
The real sample rate
session_id
- The PCM byte rate from the logging helper.
32.0is correct. - The
source,user_idand HTTP status of the session request. Never thesession_tokenor your API key. - Whether the text message test got a reply.