Skip to main content
Most broken integrations fail the same way. The WebSocket opens, session_created arrives, audio goes out, and nothing comes back. There is no error. This happens because the server cannot check your audio. Raw audio has no header. So 48 kHz audio looks the same as 16 kHz audio. The server accepts it, and the agent never hears anything it understands.
Try this first. Send one text message on the session:
  • The agent replies? Your token, socket, agent and billing all work. The problem is in your audio code.
  • No reply? The problem is not your audio. Check the steps below from the top.

Add logging first

Attach this helper right after new WebSocket(...). It logs session_created, every error, the close code, and how many bytes of audio you send per second.
A correct voice session logs 32.0 kB/s PCM while the microphone is on. Any other number points to the bug.

The agent never replies

Go through these checks in order.
1

Did session_created arrive?

The server sends session_created first, once, on every session.
  • It never arrived. The server closed the socket during setup. Read event.code in onclose and look it up in Close codes. Do not retry with the same token. It is already used, so you will get 4001 and hide the real cause.
  • It arrived. Go to the next step.
2

Did you get any error messages?

The server sends {"type": "error", "code": "..."} for anything it cannot use. The session stays open.
  • no_audio_received: you are sending images but no audio. Go to step 6.
  • Any other code: look it up in Error messages.
  • No errors: the server accepts your messages. The problem is inside the audio itself. Go to step 3.
text sessions do not send these errors. They ignore audio and images without saying anything.
3

Is audio being sent?

Check the byte rate in the logging helper. 16,000 samples × 2 bytes × 1 channel = 32,000 bytes per second.An AudioContext created outside a click starts suspended. It sends nothing until resume() finishes, with no error.Send audio all the time, including silence. The agent uses the silence to tell when the user has stopped talking.
4

Is the sample rate right?

This is the most common cause. new AudioContext({ sampleRate: 16000 }) is only a request. Safari and Firefox often ignore it and use 44.1 or 48 kHz. Nothing tells you.Check what you actually got:
If the rate is wrong, the agent hears your voice three times too slow and far too deep. It never detects speech. There is no error anywhere.Two fixes: use the client library, which resamples for you, or resample yourself:
Do not resample by dropping samples. It distorts the audio. The agent may then hear the wrong words, which is harder to spot than silence.
5

Is the format right?

The format is raw PCM: 16 kHz, 16-bit signed little-endian, mono, no WAV header, base64-encoded. See Media format.Four mistakes all look like a silent agent:
  • A WAV header. You are probably encoding whole files instead of streaming.
  • Float32 samples. The byte rate shows 64.0 kB/s at 16 kHz.
  • Big-endian samples. Only happens if you build bytes with DataView.setInt16 and leave out littleEndian. It defaults to big-endian.
  • Stereo. Two channels read as one fast, garbled channel.
Convert Float32 to Int16 like this:
Send small chunks as you record. 20 to 100 ms per message works well. Do not send a whole sentence in one message.
6

Camera or screen: are you sending microphone audio?

The agent only replies to speech. Images alone never get a reply. After about 8 seconds of images with no audio, the server sends no_audio_received.For screen share, the audio from getDisplayMedia is the computer’s sound, not the microphone. Get the two separately and merge them:
7

Does a text message get a reply?

A text/plain message makes the agent reply right away, with no audio involved. It works on every session type.
Send text/plain data as plain text, not base64. If you base64 it, the agent reads the code out loud.You can also test with no browser at all. Create a session:
Then run this script with the ws_url within 60 seconds:
If the script gets a reply and your browser does not, the problem is in your browser code. It is not your account, agent or plan.

Other problems

Your app must stop the agent’s audio. When the user starts talking, clear the playback queue right away. Otherwise the agent keeps playing everything already buffered.Detect speech locally on the microphone, and clear the queue when speech starts. The reference code uses an energy threshold of 0.015, 0.6 seconds of speech to start, and 0.8 seconds of silence to stop.The server also sends interrupted: true. It arrives later, so use it as a backup. See Barge-in.With these settings, the agent keeps talking for about 600 ms after the user starts. A very short word, like “wait”, may not interrupt it. You can lower speechSeconds to react faster. Background noise will then interrupt the agent more often.
The agent’s audio is continuous. Gaps come from how you play it.
  • One AudioBufferSourceNode per chunk adds a gap between chunks. Use one continuous queue instead.
  • Clicks between chunks mean chunks are lost or out of order. Play them in the order they arrive.
  • Clicks all the time mean clipping. Clamp Float32 samples to [-1, 1] before you convert.
Play audio through a separate AudioContext at 24000 Hz.
The microphone picks up the agent’s voice from the speakers. The agent thinks the user is talking and stops.
  • Turn on echo cancellation: getUserMedia({ audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true } }).
  • Test with headphones. If the problem goes away, it is echo.
  • Check input_transcription. If it contains the agent’s own words, this is the cause.
Speed and pitch problems always mean a sample rate mismatch.
  • Agent too fast and high: you play 24 kHz audio in a faster context without adjusting. Create the playback context at 24000.
  • Agent too slow and deep: you play 24 kHz audio as if it were 16 kHz.
  • Agent mishears you: your input audio is at the wrong rate or distorted. See steps 4 and 5.
  • Both wrong: you use one AudioContext for recording and playback. Use two.
Check these in order:
  1. Wrong source. The default is audio. An audio session tells the agent it cannot see anything, even if you send images. Create the session with "source": "screen" or "camera".
  2. No microphone audio. No turn starts. See step 6.
  3. The screen just changed. The agent gets at most one image per second. Wait a moment after switching tabs.
  4. Blank images. Drawing a video before it has loaded gives a black image. Check video.videoWidth && video.readyState >= 2 before drawing.
  5. Images too small. Use a longest side of 1280 px and quality 0.8.
  6. The user stopped sharing. The browser’s Stop sharing button ends the video track but not the socket. Listen for the track’s ended event.
There are two separate causes.The browser cannot create sessions. The API accepts browser requests from http://localhost:3000, but not from your domain. So creating sessions from the browser works locally and fails with a CORS error when deployed. Create sessions on your server. See Authentication.The page is not HTTPS. The microphone and screen capture only work on secure pages. On a deployed http:// page, navigator.mediaDevices is undefined. Serve over HTTPS.
You can create 60 sessions per hour and 500 per day per workspace. Each connection attempt uses one. These use them up fast:
  • Creating a session when a component mounts. React Strict Mode doubles this in development.
  • Creating a session on every hot reload.
  • Reconnecting in a loop after every close.
Create sessions on a user action, like a button click. Add a limit and a delay to automatic reconnects.
The socket closed with 4001. Usually the token was already used, often by a double-run effect, a hot reload, or a retry. It may also have expired after 60 seconds. See 4001: invalid token.
Read the close reason.
  • “Too many connections”, or 1006 with no reason: too many attempts from your IP. Wait a minute.
  • “Too many concurrent sessions for your plan”: old sessions still hold slots. Always call stop() on beforeunload. A session closed without cleanup can hold its slot for up to 2 hours.
Use a new user_id for each session, and reuse it only for retries of that same session.
  • Not HTTPS. navigator.mediaDevices does not exist on insecure pages.
  • No user action. getDisplayMedia must be called directly in a click handler.
  • Blocked before. If the user blocked the microphone once, the browser remembers. They must reset it in site settings.
  • Inside an iframe. The iframe needs allow="microphone; camera; display-capture".
  • Permissions-Policy. A strict Permissions-Policy header on your page blocks it.
  • No device. With no microphone, you get NotFoundError.
Log the error’s name. NotAllowedError, NotFoundError, InvalidStateError and TypeError each point to a different cause.

Known limits

These are not bugs in your code.
  • You cannot fetch a web session through GET /api/v1/calls/{call_id}. Collect the transcript in the browser, or use a webhook.
  • A dropped session cannot be resumed. Create a new session and connect again.
  • No keepalive is needed. There is no idle timeout or maximum length to handle. If sessions end after the same time every time, check for a proxy or load balancer timeout on your side.

Still stuck?

Collect these before you ask for help:

Close code and reason

event.code and event.reason from onclose.

The first error

The first error message, with its code and text.

The real sample rate

context.sampleRate, plus the browser and version.

session_id

From session_created. It links your session to our logs.
Also include:
  • The PCM byte rate from the logging helper. 32.0 is correct.
  • The source, user_id and HTTP status of the session request. Never the session_token or your API key.
  • Whether the text message test got a reply.
This helper collects most of it for you: