Skip to main content
A web session is a live conversation between a browser and an agent. It runs over one WebSocket connection.
  • The browser sends the user’s microphone audio. In camera and screen sessions, it also sends still images.
  • The agent sends back its voice, live transcripts of both sides, and any tool calls it makes.
Turn-taking is automatic. The agent listens to the audio and decides when the user has finished speaking. There is no “send” button.

Choose how to build

All three support voice, camera and screen share. Start with the widget if you can. Move to the library or the protocol only when you need more control.

How a session starts

The widget does all of this for you. If you use the library or the protocol, every session follows these steps:
  1. The browser asks your server for a session. This is a normal request to your own backend. Check here that the user is allowed to talk to the agent.
  2. Your server creates the session. It calls POST /api/v1/realtime/sessions with your API key. The key never leaves your server.
  3. Eclatira checks the request. It checks billing, the agent, the session type and rate limits. If any check fails, you get an HTTP error here, before any connection opens.
  4. Your server returns ws_url to the browser.
  5. The browser opens ws_url. Audio goes straight between the browser and Eclatira. Your server is not in the media path.
Your server must create the session. The browser cannot do it, even in a demo. See Keep your key on the server.

Create a session

string
required
The agent to talk to.
string
required
A unique ID for this session. Generate a new one for each session, such as a UUID. See the rules below.
string
default:"audio"
The session type: audio, camera, screen or text.
object
Values for the {{placeholders}} in the agent’s instructions and first message. Keys and values are strings.

Session types

You choose the type with source when you create the session. You cannot change it later. To switch, end the session and start a new one. Camera and screen sessions send images the same way. Both use image/jpeg on the same socket as the audio. The type only changes what the agent is told it can see.

Rules every integration must follow

The session_token expires 60 seconds after you create it. It is used up by the first connection attempt, even if that attempt fails.Create a new session for every connection, including reconnects. A dropped connection cannot be resumed.
The server uses user_id as the session’s identity. If two sessions are open at the same time with the same user_id, they break each other.This usually happens when you use your app’s own user ID and the user opens two tabs. Generate a new random value for each session. Keep your own mapping to your real user if you need it.Reusing a user_id after the old session has closed is safe. Voice, camera and screen sessions always start fresh. They never carry over an earlier conversation.
ws_url is complete. Do not add, remove or change any part of it.
The agent only replies to speech. A session that sends images and no audio connects fine and then waits forever.To make the agent respond without speech, send a text message. In the library, call session.sendText("What do you see?").
You can create 60 sessions per hour and 500 per day per workspace. Over the limit, you get 429 with a Retry-After header.Every connection attempt uses one. A debug loop that reconnects on every error uses them up fast.

Before you start

1

An agent

Create one in the dashboard or with POST /api/v1/agents. Camera and screen are on by default.
2

An API key

Create one under Settings → API Keys. See Authentication.
3

A server

Any backend that can keep a secret and answer one HTTP request. It creates sessions for the browser.

Save the transcript

You have two ways to keep a record of a web session:
  • In the browser. Collect input_transcription (the user) and output_transcription (the agent) as they arrive. Send them to your own backend as you go.
  • With a webhook. A registered webhook gets call.completed with the full transcript when the session ends.
You cannot fetch a web session later with GET /api/v1/calls/{call_id}. That endpoint only finds phone calls.

Next steps

Embed widget

Voice, camera and screen share with two script tags.

Voice in the browser

Build your own UI with the client library.

WebSocket protocol

Every message the socket sends and receives.

Troubleshooting

The agent connects but never speaks.