Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -56,3 +56,14 @@ jobs:
cargo build --locked --all-features
cargo test --locked --all-features
echo "::endgroup::"

# The `livekit_native_*` tests construct real libwebrtc `NativeAudioSource`
# objects. libwebrtc's global runtime intermittently segfaults under the
# heavy thread concurrency of the full unit-test binary, so those tests are
# `#[ignore]`d out of the run above and executed here in their own isolated,
# single-threaded process instead.
- name: Test LiveKit native audio (isolated)
run: |
echo "::group::Running libwebrtc-native loading-audio tests in isolation"
cargo test --locked --all-features --lib -- --ignored --test-threads=1 livekit_native_
echo "::endgroup::"
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -11,3 +11,6 @@ libonnxruntime.so
onnxruntime/
dist/
.mcgravity/

# Local planning doc for loading-indicator feature (not part of the product tree)
loading-indicator.md
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,12 @@

All notable changes to Sayna will be documented in this file.

## [Unreleased]

### Features

- **websocket:** Add loading indicator audio (`loading_audio` in `config`, `loading_start` / `loading_stop` commands) on a dedicated `"loading-audio"` LiveKit track, independent of TTS

## [0.1.15] - 2026-04-30

### Features
Expand Down
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -225,9 +225,10 @@ See existing implementations for patterns:

### WebSocket
- `/ws` - Real-time voice processing
- Receives: Config, audio data, speak commands, **clear command**, send_message, sip_transfer
- Receives: Config (optionally with `loading_audio`), audio data, speak commands, **clear command**, loading_start, loading_stop, send_message, sip_transfer
- Sends: Ready, STT results, TTS audio, messages, participant_connected, participant_disconnected, track_subscribed, tts_playback_complete, vad_event, error, sip_transfer_error
- **Clear command**: Immediately stops TTS and clears audio buffers (fire-and-forget, respects `allow_interruption` setting)
- **Loading commands**: `loading_start`/`loading_stop` control a looping loading-indicator audio clip on a dedicated `"loading-audio"` LiveKit track (fire-and-forget, independent of `speak`/`clear`)

### Webhooks
- `POST /livekit/webhook` - LiveKit event receiver (signature verified)
Expand Down
31 changes: 30 additions & 1 deletion docs/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -544,7 +544,7 @@ This enables efficient trunk management without manual provisioning.
1. Connect via WebSocket and immediately send a `config` message.
2. The server initializes providers (and LiveKit, if requested) and replies with `ready`.
3. Stream audio frames as binary messages that match the declared STT sample rate/encoding.
4. Use `speak`, `clear`, or `send_message` commands to drive TTS and LiveKit data.
4. Use `speak`, `clear`, `loading_start`, `loading_stop`, or `send_message` commands to drive TTS, the loading indicator, and LiveKit data.
5. Close the socket when finished; the server also closes when a fatal `error` is emitted.

#### Incoming Messages
Expand All @@ -560,6 +560,7 @@ Configures audio processing and optional LiveKit mirroring. Must be the first me
| `stt_config` | object | Conditional | Required when `audio=true`. See table below. |
| `tts_config` | object | Conditional | Required when `audio=true`. Same schema as the REST `tts_config`. |
| `livekit` | object | No | LiveKit options; omitted for WebSocket-only sessions. |
| `loading_audio` | object | No | Optional loading indicator clip. Processed only when `audio=true` and `livekit` is present; a decode failure is non-fatal. See table below. |

**STT configuration**

Expand All @@ -583,6 +584,20 @@ Configures audio processing and optional LiveKit mirroring. Must be the first me
| `sayna_participant_name` | string | Override for the agent display name (default `Sayna AI`). |
| `listen_participants` | array<string> | Restrict audio/data processing to specific participant identities (empty list listens to all). |

**`loading_audio` configuration**

Decoded once at config time; the validated clip is held for the session and played on a dedicated `"loading-audio"` LiveKit track via the `loading_start` / `loading_stop` commands. Only 16-bit PCM is supported.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `data` | string | Yes | Base64-encoded audio bytes: a complete WAV file or raw PCM. |
| `format` | string | No | `wav` or `pcm`. Auto-detected from the `RIFF`/`WAVE` signature when omitted. |
| `sample_rate` | integer | Conditional | Sample rate in Hz (8000–48000). Required for raw PCM; ignored for WAV. |
| `channels` | integer | No | Channel count for raw PCM (`1` or `2`, default `1`); ignored for WAV. |
| `volume` | number | No | Playback volume `0.0`–`1.0` (default `1.0`); out-of-range values are clamped. |

See [Loading Indicator](websocket.md#loading-indicator) in the WebSocket reference for authoring guidance and limits.

##### `speak`
Queues text for synthesis.

Expand All @@ -600,6 +615,20 @@ Immediately clears queued audio and LiveKit buffers. Useful for interruptions.
| --- | --- | --- |
| `type` | string | `clear`. |

##### `loading_start`
Begins looping the configured loading indicator clip on a dedicated `"loading-audio"` LiveKit track. Requires `audio=true`, an active LiveKit room, and a `loading_audio` clip supplied in `config`. Fire-and-forget on success; idempotent if the loop is already running; failures emit an `error` message.

| Field | Type | Description |
| --- | --- | --- |
| `type` | string | `loading_start`. |

##### `loading_stop`
Stops the loading indicator loop with a short fade-out. Fire-and-forget; idempotent and always silent — stopping when nothing is playing is never an error. The loop is controlled only by `loading_start` / `loading_stop`; `speak` and `clear` do not affect it.

| Field | Type | Description |
| --- | --- | --- |
| `type` | string | `loading_stop`. |

##### `send_message`
Publishes a LiveKit data message (if LiveKit is configured) and also feeds the VoiceManager message bus.

Expand Down
36 changes: 27 additions & 9 deletions docs/livekit_integration.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,15 +17,16 @@ Sayna provides comprehensive LiveKit integration with automatic SIP infrastructu
## Table of Contents

1. [Architecture Overview](#architecture-overview)
2. [Configuration](#configuration)
3. [Inbound Webhooks (LiveKit → Sayna)](#inbound-webhooks-livekit--sayna)
4. [SIP Configuration & Auto-Provisioning](#sip-configuration--auto-provisioning)
5. [Outbound SIP Calls](#outbound-sip-calls)
6. [Outbound Webhooks (Sayna → Downstream)](#outbound-webhooks-sayna--downstream)
7. [Webhook Signing](#webhook-signing)
8. [Testing & Development](#testing--development)
9. [Operations & Troubleshooting](#operations--troubleshooting)
10. [Security Considerations](#security-considerations)
2. [Published Audio Tracks](#published-audio-tracks)
3. [Configuration](#configuration)
4. [Inbound Webhooks (LiveKit → Sayna)](#inbound-webhooks-livekit--sayna)
5. [SIP Configuration & Auto-Provisioning](#sip-configuration--auto-provisioning)
6. [Outbound SIP Calls](#outbound-sip-calls)
7. [Outbound Webhooks (Sayna → Downstream)](#outbound-webhooks-sayna--downstream)
8. [Webhook Signing](#webhook-signing)
9. [Testing & Development](#testing--development)
10. [Operations & Troubleshooting](#operations--troubleshooting)
11. [Security Considerations](#security-considerations)

---

Expand Down Expand Up @@ -91,6 +92,23 @@ Sayna provides comprehensive LiveKit integration with automatic SIP infrastructu

---

## Published Audio Tracks

When a WebSocket session is audio-enabled and joins a LiveKit room, the Sayna agent participant publishes synthesized speech on an audio track named `"tts-audio"`.

If that session's `config` message also includes a `loading_audio` object (see [Loading Indicator](websocket.md#loading-indicator) in the WebSocket reference), the agent participant also publishes a **second** audio track named `"loading-audio"`, carrying the loading indicator sound. The agent then publishes **two** audio tracks:

| Track name | Carries |
|------------|---------|
| `"tts-audio"` | Synthesized speech (text-to-speech output). |
| `"loading-audio"` | The loading indicator clip, looped while the calling application is busy. |

The two tracks are independent and can be audible at the same time. LiveKit client SDKs automatically play every subscribed audio track, so participants hear both without any extra client-side handling.

**Recordings:** Because `"loading-audio"` is a real published track, it is included in room-composite egress recordings alongside `"tts-audio"`. This is correct behavior — the recording faithfully reflects what the human participant heard.

---

## Configuration

### Required Environment Variables
Expand Down
61 changes: 61 additions & 0 deletions docs/openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -661,6 +661,11 @@ components:
- type: 'null'
- $ref: '#/components/schemas/LiveKitWebSocketConfig'
description: Optional LiveKit configuration for real-time audio streaming
loading_audio:
oneOf:
- type: 'null'
- $ref: '#/components/schemas/LoadingAudioConfig'
description: Optional loading-indicator audio configuration.
stream_id:
type:
- string
Expand Down Expand Up @@ -791,6 +796,24 @@ components:
- type: 'null'
- $ref: '#/components/schemas/VADConfigUpdate'
description: VAD configuration updates (silence threshold, etc.)
- type: object
description: Begin looping the configured loading-indicator audio into the LiveKit room.
required:
- type
properties:
type:
type: string
enum:
- loading_start
- type: object
description: Stop the loading-indicator audio loop (with a short fade-out).
required:
- type
properties:
type:
type: string
enum:
- loading_stop
description: WebSocket message types for incoming messages
ListRoomsResponse:
type: object
Expand Down Expand Up @@ -858,6 +881,44 @@ components:
- 'null'
description: Sayna AI participant display name (defaults to "Sayna AI")
example: Sayna AI
LoadingAudioConfig:
type: object
description: Loading-indicator audio configuration supplied in the `config` message.
required:
- data
properties:
channels:
type:
- integer
- 'null'
format: int32
description: Channel count for raw PCM (1 = mono, 2 = stereo); defaults to 1. Ignored for WAV.
example: 1
minimum: 0
data:
type: string
description: Base64-encoded audio bytes — a complete WAV file or raw 16-bit PCM.
format:
type:
- string
- 'null'
description: 'Audio format: "wav" or "pcm". If omitted, the server auto-detects.'
example: wav
sample_rate:
type:
- integer
- 'null'
format: int32
description: Sample rate in Hz. Required for raw PCM; ignored for WAV (header is authoritative).
example: 16000
minimum: 0
volume:
type:
- number
- 'null'
format: float
description: Playback volume from 0.0 (silent) to 1.0 (authored level). Default 1.0; out-of-range clamped.
example: 0.3
MuteParticipantRequest:
type: object
description: |-
Expand Down
Loading