Skip to main content
This page is the byte-level spec for a web session. Use it if you write your own client instead of using the library. If a session connects and the agent never speaks, the cause is almost always on this page.
Most format mistakes fail with no error. The server does not inspect your audio. It passes the bytes to the agent as 16 kHz, 16-bit, mono PCM, whatever they really are. A wrong sample rate, bit depth, byte order, channel count or a WAV header all look the same: the socket stays open, no error arrives, and the agent never replies.

Summary

The input and output rates are different. You cannot change either one. The agent only takes a turn when it hears speech in the audio. There is no end-of-turn message. Images never start a turn.

Input audio

Required format

Four samples with values 0, 1000, -1000, 32767 look like this in bytes:
The byte count is always samples × 2. There is no minimum length and no alignment rule.

Convert Float32 to Int16

Web Audio gives you samples from −1 to 1. Scale negative and positive values differently, because the Int16 range is not symmetric:
Always clamp first. Web Audio can produce values outside −1 to 1. Without the clamp, those values wrap around and cause clicks.

The AudioContext sample-rate problem

This is the most common reason a correct-looking client never gets a reply.
sampleRate is a request. The browser can give you a different rate, and some do: Nothing warns you. You send 48 kHz audio labeled as 16 kHz. The agent hears one second of speech as three seconds of slow, deep sound. It never detects speech. Always check the rate you got:
Do not check the browser name. Read context.sampleRate and resample from that. The resampler below costs nothing when the rate already matches.

The resampling worklet

This is the same recorder the client library uses. It reads the real rate, resamples to 16 kHz, converts to Int16, and posts the bytes to the main thread. It also detects speech locally for barge-in. Save it as eclatira-recorder.js and serve it from your own site.
Wire it up on the main thread:
The library includes this worklet inside its own file and loads it from a Blob URL. You do not need to host anything if you use the library.

Why the resampler keeps state

process() runs once per block of 128 samples. The resampler steps through the input by a fraction, so two values must carry over from one block to the next: These bugs make a hand-written resampler sound almost right, while speech detection fails.

How often to send

Send audio all the time, including silence. The agent needs the silence to know the user stopped talking. If you only send while the user speaks, the agent either interrupts all the time or never answers. To make the agent reply without speech, send text. This starts a turn right away:
text/plain data is plain UTF-8, not base64. It is the only type that is not base64.

Output audio

The agent’s voice arrives as parts entries:
Chunk sizes vary. Do not assume a fixed size.

You need two AudioContexts

An AudioContext has one sample rate. Recording needs 16000 and playback needs 24000. Create two:
With one context, one side is always wrong. At 16000, the agent sounds slow and low. At 24000, your microphone audio goes out at the wrong rate and the agent never replies.

The playback worklet

Playing each chunk with its own BufferSource adds gaps. It also gives you nothing to clear on barge-in. Use a queue in a worklet. Save it as eclatira-player.js:
Wire it up and decode incoming audio:
Dividing by 32768 on playback is correct. Dividing by 32767 would let -32768 go slightly past −1 and clip.

Barge-in

When the user talks over the agent, the agent’s audio must stop at once. If it does not, both voices overlap and the conversation falls apart.

Why the server signal is not enough

The server sends interrupted: true when it detects the user cutting in. Handle it, but do not rely on it alone: Your playback queue already holds all the audio the server sent. If you do not clear it, the agent keeps talking until the queue is empty. Muting or ignoring new chunks does not help.

Detection settings

These values are in the recorder worklet above: Lower values make the agent stop on room noise. Higher values make interruptions feel slow. Test changes with a real microphone in a real room, one at a time.

Minimum delay

With speechSeconds at 0.6, the queue cannot clear sooner than about 600 ms after the user starts talking. Measured, it clears at 632 ms. So:
  • About 600 ms of agent speech plays over the user even when everything works. This is expected.
  • A very short sound never clears the queue locally. Only the slower server signal catches it.
The local detector only clears playback. It does not decide turns.

Wire up both triggers

Keep sending microphone audio while you clear playback. The agent needs it to understand the interruption.

Video frames

Camera and screen sessions send images the same way, on the same socket as audio. The only difference is what the agent is told it is looking at.

Format

canvas.toDataURL("image/jpeg", q) returns data:image/jpeg;base64,.... Send only the part after the comma. The server removes the prefix if you leave it, but a value with no comma fails with invalid_base64.

Server pacing

The agent always gets the most recent image. Sending one about every 2 seconds means a fresh image is waiting when each window opens. Sending at video speed gains nothing. The server sets no limits, so these are choices, not rules:
Google’s own reference client is more careful: 0.5 images per second, scaled to 25%. Move toward those numbers for very large screens or long sessions. Eclatira’s defaults favor readable text.

Capture an image

Keep the video.readyState < 2 check. Drawing a video before it has a frame gives a black image. The server accepts it, and the agent describes a black screen.

Video with no audio

A camera or screen session with no microphone audio never gets a reply. The session connects, accepts every image, and nothing else happens.
The server sends this once, when an image arrives at least 8 seconds after the first image and no audio has arrived:
The session stays open. This check only runs when an image arrives. If you stop sending images, you never get it. So no error does not prove audio is arriving. On screen sessions, the usual cause is getDisplayMedia:

Message size limit

The only size limit is the server’s WebSocket limit: 16 MiB per message, including the base64 overhead. That means about 12 MiB of JPEG. A message over the limit closes the connection with no error message. A 1280 px JPEG at quality 0.8 is usually 100 to 400 KB, far below the limit.

The message envelope

Every message you send is a JSON object in a WebSocket text frame, with two fields:

Other accepted spellings

Write audio/pcm, image/jpeg and text/plain. The server also accepts these, so a near-miss does not fail silently:

The rate parameter is ignored

audio/pcm;rate=48000 does not make the server resample. The server removes the rate and always treats the audio as 16 kHz. Resample in your client.

Binary frames

Binary WebSocket frames are not supported. Each one gets this reply:
The session stays open, but the frame is dropped. A client that only sends binary audio never gets a reply. For every error code, see Errors and close codes.

Silent failures

Every row below leaves the WebSocket open and healthy. Unless the row says so, there is no error.

Outside the browser

With ffmpeg, the format maps to these flags:
In Python, always pack little-endian explicitly:

Next steps

WebSocket protocol

Every message the server sends.

Troubleshooting

Step-by-step checks for a silent agent.