> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eclatira.com/llms.txt
> Use this file to discover all available pages before exploring further.

# How web sessions work

> A web session is a live voice conversation with an agent in the browser. Learn how one starts, the four session types, and the rules every integration must follow.

A web session is a live conversation between a browser and an agent. It runs over one WebSocket connection.

* The browser sends the user's microphone audio. In camera and screen sessions, it also sends still images.
* The agent sends back its voice, live transcripts of both sides, and any tool calls it makes.

Turn-taking is automatic. The agent listens to the audio and decides when the user has finished speaking. There is no "send" button.

## Choose how to build

| | Widget | Library | Protocol |
| - | - | - | - |
| **You write** | Two script tags | Your UI and a token endpoint | Everything |
| **Media code** | Done for you | Done for you | You build it |
| **Custom UI** | Colors and size | Yes | Yes |
| **Live transcripts** | No | Yes | Yes |
| **Guide** | [Embed widget](/embed-widget) | [Voice in the browser](/realtime/browser-voice) | [Protocol](/realtime-protocol) |

All three support voice, camera and screen share. Start with the widget if you can. Move to the library or the protocol only when you need more control.

## How a session starts

The widget does all of this for you. If you use the library or the protocol, every session follows these steps:

```mermaid theme={"system"}
sequenceDiagram
    participant B as Browser
    participant S as Your server
    participant E as Eclatira
    B->>S: 1. Ask for a session
    S->>E: 2. POST /api/v1/realtime/sessions (with ek_ key)
    E-->>S: 3. session_token, ws_url, expires_at
    S-->>B: 4. ws_url
    B->>E: 5. Open WebSocket to ws_url
    E-->>B: session_created
    B-)E: microphone audio, images
    E-)B: agent audio, transcripts
```

1. **The browser asks your server for a session.** This is a normal request to your own backend. Check here that the user is allowed to talk to the agent.
2. **Your server creates the session.** It calls `POST /api/v1/realtime/sessions` with your API key. The key never leaves your server.
3. **Eclatira checks the request.** It checks billing, the agent, the session type and rate limits. If any check fails, you get an HTTP error here, before any connection opens.
4. **Your server returns `ws_url` to the browser.**
5. **The browser opens `ws_url`.** Audio goes straight between the browser and Eclatira. Your server is not in the media path.

<Warning>
  Your server must create the session. The browser cannot do it, even in a demo. See [Keep your key on the server](/authentication#keep-your-key-on-the-server).
</Warning>

### Create a session

```bash theme={"system"}
curl -X POST https://app.eclatira.com/api/v1/realtime/sessions \
  -H "Authorization: Bearer ek_your_api_key_here" \
  -H "Content-Type: application/json" \
  -d '{
    "agent_id": "your-agent-id",
    "user_id": "session-7f3c1e9a",
    "source": "audio"
  }'
```

```json theme={"system"}
{
  "session_token": "rt_...",
  "expires_at": "2026-09-21T10:15:00.482913+00:00",
  "ws_url": "wss://app.eclatira.com/ws/session-7f3c1e9a?session_token=rt_...&source=audio"
}
```

<ParamField body="agent_id" type="string" required>
  The agent to talk to.
</ParamField>

<ParamField body="user_id" type="string" required>
  A unique ID for this session. Generate a new one for each session, such as a UUID. See the rules below.
</ParamField>

<ParamField body="source" type="string" default="audio">
  The session type: `audio`, `camera`, `screen` or `text`.
</ParamField>

<ParamField body="variables" type="object">
  Values for the `{{placeholders}}` in the agent's instructions and first message. Keys and values are strings.
</ParamField>

## Session types

You choose the type with `source` when you create the session. You cannot change it later. To switch, end the session and start a new one.

| `source` | The browser sends | Notes |
| - | - | - |
| `audio` | Microphone audio | The default. |
| `camera` | Microphone audio **and** webcam images | Needs `enable_webcam` on the agent. It is on by default. |
| `screen` | Microphone audio **and** screen images | Needs `enable_screen_share` on the agent. It is on by default. |
| `text` | Typed messages | No audio. Uses different server messages. |

Camera and screen sessions send images the same way. Both use `image/jpeg` on the same socket as the audio. The type only changes what the agent is told it can see.

## Rules every integration must follow

<AccordionGroup>
  <Accordion title="Each token works once, for 60 seconds" icon="clock">
    The `session_token` expires 60 seconds after you create it. It is used up by the first connection attempt, even if that attempt fails.

    Create a new session for every connection, including reconnects. A dropped connection cannot be resumed.
  </Accordion>

  <Accordion title="Use a new user_id for every session" icon="fingerprint">
    The server uses `user_id` as the session's identity. If two sessions are open at the same time with the same `user_id`, they break each other.

    This usually happens when you use your app's own user ID and the user opens two tabs. Generate a new random value for each session. Keep your own mapping to your real user if you need it.

    Reusing a `user_id` after the old session has closed is safe. Voice, camera and screen sessions always start fresh. They never carry over an earlier conversation.
  </Accordion>

  <Accordion title="Open ws_url exactly as returned" icon="link">
    `ws_url` is complete. Do not add, remove or change any part of it.
  </Accordion>

  <Accordion title="Camera and screen sessions also need the microphone" icon="microphone">
    The agent only replies to speech. A session that sends images and no audio connects fine and then waits forever.

    To make the agent respond without speech, send a text message. In the library, call `session.sendText("What do you see?")`.
  </Accordion>

  <Accordion title="Rate limits" icon="gauge">
    You can create **60 sessions per hour** and **500 per day** per workspace. Over the limit, you get `429` with a `Retry-After` header.

    Every connection attempt uses one. A debug loop that reconnects on every error uses them up fast.
  </Accordion>
</AccordionGroup>

## Before you start

<Steps>
  <Step title="An agent">
    Create one in the dashboard or with `POST /api/v1/agents`. Camera and screen are on by default.
  </Step>

  <Step title="An API key">
    Create one under **Settings → API Keys**. See [Authentication](/authentication).
  </Step>

  <Step title="A server">
    Any backend that can keep a secret and answer one HTTP request. It creates sessions for the browser.
  </Step>
</Steps>

## Save the transcript

You have two ways to keep a record of a web session:

* **In the browser.** Collect `input_transcription` (the user) and `output_transcription` (the agent) as they arrive. Send them to your own backend as you go.
* **With a webhook.** A registered [webhook](/webhooks) gets `call.completed` with the full transcript when the session ends.

You cannot fetch a web session later with `GET /api/v1/calls/{call_id}`. That endpoint only finds phone calls.

## Next steps

<CardGroup cols={2}>
  <Card title="Embed widget" icon="window-maximize" href="/embed-widget">
    Voice, camera and screen share with two script tags.
  </Card>

  <Card title="Voice in the browser" icon="microphone" href="/realtime/browser-voice">
    Build your own UI with the client library.
  </Card>

  <Card title="WebSocket protocol" icon="plug" href="/realtime-protocol">
    Every message the socket sends and receives.
  </Card>

  <Card title="Troubleshooting" icon="bug" href="/realtime/troubleshooting">
    The agent connects but never speaks.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.