Skip to main content
This guide builds a working voice conversation in a browser tab. The user can interrupt the agent at any time. Every code sample is a complete file you can copy. You need three parts:
  1. A token endpoint on your server. It creates a session with your API key.
  2. The client library, @eclatira/realtime, on your page.
  3. A page with a start button.
The library handles the hard parts of browser audio. It records the microphone at the right sample rate, plays the agent’s voice without gaps, and stops playback when the user interrupts. You write no audio code.
1

Create an agent

Skip this step if you already have an agent. You only need its agent_id.
The response includes the agent_id:
Camera and screen share are on by default, so the same agent also works for camera and screen sessions.
2

Add a token endpoint to your server

This is the only server code you need. It keeps the API key, creates one session, and returns the response to the browser.
The endpoint returns this to the browser:
Put your own login check where the comment says. Anyone who can call this endpoint can start a session that bills your account.
Read the session rules before you go live. The short version: each token works once, for 60 seconds, and each session needs a new user_id.
3

Install the library

@eclatira/realtime is one ES module with no dependencies and no build step.
4

Build the page

Save this file. Serve it from http://localhost or HTTPS. The microphone does not work on file:// pages. Set TOKEN_URL to your endpoint from step 2.
voice.html
Open the page and click Start. Allow the microphone and start talking.
Call session.start() from a click or other user action. Browsers only allow the microphone and audio playback after one.
If your page and your token endpoint are on different origins, set ALLOWED_ORIGIN to the page’s exact origin. For example, the page is on :3000 and the Node server is on :8787.The library sends the request with credentials: "include". Browsers reject Access-Control-Allow-Origin: * on these requests. The symptom is a CORS error and mint_failed from session.start(). A same-origin route, like the Next.js example, avoids this.

Know when the session is live

The WebSocket opening does not mean the session started. When the server refuses a session, it accepts the socket first and then closes it with a code. So onopen fires even for an expired token or an unpaid account. The real signal is the first session_created message. session.start() waits for it:
  • It resolves when session_created arrives.
  • It rejects with an EclatiraError if the socket closes first. The error’s code names the reason: invalid_token, rate_limited, limit_reached, billing_blocked, source_not_enabled or server_error.
If you use the raw protocol, do the same. Only show “connected” after session_created.

Library reference

Constructor options

string
Your token endpoint. It must return { ws_url }. The library adds ?source=... to the URL. Required unless you pass getSession.
async function
Use this instead of tokenUrl when a plain GET is not enough. For example, when you need a POST body or your own auth header. It must return { ws_url }.
number
default:"2000"
How often to send an image, in milliseconds. Camera and screen sessions only.
number
default:"0.8"
JPEG quality of each image. Camera and screen sessions only.
number
default:"1280"
Longest side of each image, in pixels. Camera and screen sessions only.

Methods

Events

agentaudio gives you the size of each audio chunk in bytes. The agent’s audio is 24 kHz, 16-bit, mono. So bytes / 48000 is the chunk’s length in seconds.

Transcripts

transcript events send the full text so far for the current turn, not just the new words. Replace the current line each time. Start a new line when final is true. The example page does this in write(). To keep a record, send each final transcript to your server as it arrives. Or register a webhook. It gets the full transcript when the session ends.

End the session

Call session.stop() from your stop button. Also call it on beforeunload, as the example does. Your plan limits how many sessions can run at the same time. A session releases its slot when it disconnects. If you reload a page that never called stop(), old sessions can hold slots for a while. New connections then close with code 4029.
If you keep getting 4029 while you develop, wait for old sessions to drop. Do not keep retrying. Each attempt also uses one of your 60 sessions per hour.
The server has no idle timeout and no maximum session length that you need to handle. You do not need to send a keepalive or ping. In a very long call, the agent keeps its instructions but forgets the oldest part of the conversation.

Build it without the library

If you cannot use the library, you must build the audio pipeline yourself. Five things must be right. Four of them fail with no error:
  1. Record audio at 16 kHz, 16-bit signed little-endian, mono, with no WAV header. Base64 the raw bytes.
  2. Play the agent’s audio through a second AudioContext at 24 kHz.
  3. Send audio all the time, including silence. There is no end-of-turn message.
  4. Detect the user’s speech locally, and clear the playback queue the moment they start talking.
  5. Send JSON text frames only. Binary frames are rejected.
Media format explains each rule in detail. Here is a complete page that follows them:
voice-raw.html
Notes
  • The recorder node is not connected to destination. Connecting it would play the user’s microphone back through the speakers.
  • There is no required chunk size. This code sends one message per 128-sample block, about every 8 ms at 16 kHz. Messages of 20 to 100 ms also work and cost less overhead.
  • The local speech detector only clears playback. It does not decide turns. The agent does that from the audio.
  • The detector needs 0.6 seconds of speech before it fires. So the agent keeps talking for about 600 ms after the user starts. Shorter sounds, like a cough, do not interrupt the agent. You can lower speechSeconds for faster interruptions, but background noise will then interrupt the agent more often.
  • The server also sends interrupted: true. It arrives later than the local detector, so use it as a backup.
  • Agent audio chunks are continuous. Play them back to back, with no gaps.

Next steps

Camera

Let the agent see the user’s webcam.

Screen share

Let the agent see the user’s screen.

WebSocket protocol

Every message on the socket, field by field.

Errors and close codes

What each code means and how to fix it.