Summary
The input and output rates are different. You cannot change either one.
The agent only takes a turn when it hears speech in the audio. There is no end-of-turn message. Images never start a turn.
Input audio
Required format
Four samples with values
0, 1000, -1000, 32767 look like this in bytes:
samples × 2. There is no minimum length and no alignment rule.
Convert Float32 to Int16
Web Audio gives you samples from −1 to 1. Scale negative and positive values differently, because the Int16 range is not symmetric:The AudioContext sample-rate problem
This is the most common reason a correct-looking client never gets a reply.sampleRate is a request. The browser can give you a different rate, and some do:
Nothing warns you. You send 48 kHz audio labeled as 16 kHz. The agent hears one second of speech as three seconds of slow, deep sound. It never detects speech.
Always check the rate you got:
context.sampleRate and resample from that. The resampler below costs nothing when the rate already matches.
The resampling worklet
This is the same recorder the client library uses. It reads the real rate, resamples to 16 kHz, converts to Int16, and posts the bytes to the main thread. It also detects speech locally for barge-in. Save it aseclatira-recorder.js and serve it from your own site.
The library includes this worklet inside its own file and loads it from a
Blob URL. You do not need to host anything if you use the library.Why the resampler keeps state
process() runs once per block of 128 samples. The resampler steps through the input by a fraction, so two values must carry over from one block to the next:
These bugs make a hand-written resampler sound almost right, while speech detection fails.
How often to send
Send audio all the time, including silence. The agent needs the silence to know the user stopped talking. If you only send while the user speaks, the agent either interrupts all the time or never answers.
To make the agent reply without speech, send text. This starts a turn right away:
text/plain data is plain UTF-8, not base64. It is the only type that is not base64.
Output audio
The agent’s voice arrives asparts entries:
Chunk sizes vary. Do not assume a fixed size.
You need two AudioContexts
AnAudioContext has one sample rate. Recording needs 16000 and playback needs 24000. Create two:
The playback worklet
Playing each chunk with its ownBufferSource adds gaps. It also gives you nothing to clear on barge-in. Use a queue in a worklet.
Save it as eclatira-player.js:
32768 on playback is correct. Dividing by 32767 would let -32768 go slightly past −1 and clip.
Barge-in
When the user talks over the agent, the agent’s audio must stop at once. If it does not, both voices overlap and the conversation falls apart.Why the server signal is not enough
The server sendsinterrupted: true when it detects the user cutting in. Handle it, but do not rely on it alone:
Your playback queue already holds all the audio the server sent. If you do not clear it, the agent keeps talking until the queue is empty. Muting or ignoring new chunks does not help.
Detection settings
These values are in the recorder worklet above:
Lower values make the agent stop on room noise. Higher values make interruptions feel slow. Test changes with a real microphone in a real room, one at a time.
Minimum delay
WithspeechSeconds at 0.6, the queue cannot clear sooner than about 600 ms after the user starts talking. Measured, it clears at 632 ms. So:
- About 600 ms of agent speech plays over the user even when everything works. This is expected.
- A very short sound never clears the queue locally. Only the slower server signal catches it.
Wire up both triggers
Video frames
Camera and screen sessions send images the same way, on the same socket as audio. The only difference is what the agent is told it is looking at.Format
canvas.toDataURL("image/jpeg", q) returns data:image/jpeg;base64,.... Send only the part after the comma. The server removes the prefix if you leave it, but a value with no comma fails with invalid_base64.
Server pacing
The agent always gets the most recent image. Sending one about every 2 seconds means a fresh image is waiting when each window opens. Sending at video speed gains nothing.
Recommended settings
The server sets no limits, so these are choices, not rules:Google’s own reference client is more careful: 0.5 images per second, scaled to 25%. Move toward those numbers for very large screens or long sessions. Eclatira’s defaults favor readable text.
Capture an image
video.readyState < 2 check. Drawing a video before it has a frame gives a black image. The server accepts it, and the agent describes a black screen.
Video with no audio
The server sends this once, when an image arrives at least 8 seconds after the first image and no audio has arrived:getDisplayMedia:
Message size limit
The only size limit is the server’s WebSocket limit: 16 MiB per message, including the base64 overhead. That means about 12 MiB of JPEG. A message over the limit closes the connection with no error message. A 1280 px JPEG at quality 0.8 is usually 100 to 400 KB, far below the limit.The message envelope
Every message you send is a JSON object in a WebSocket text frame, with two fields:Other accepted spellings
Writeaudio/pcm, image/jpeg and text/plain. The server also accepts these, so a near-miss does not fail silently:
The rate parameter is ignored
Binary frames
Binary WebSocket frames are not supported. Each one gets this reply:Silent failures
Every row below leaves the WebSocket open and healthy. Unless the row says so, there is no error.Outside the browser
Withffmpeg, the format maps to these flags: