Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 3 additions & 4 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,11 +2,12 @@

All notable changes to Sayna will be documented in this file.

## [0.1.16] - 2026-05-25
## [Unreleased]

### Features

- **ws:** Loading-indicator audio loop on a dedicated LiveKit track by @tigranbs in [#18](https://github.com/saynaai/sayna/pull/18)
- **websocket:** Add loading indicator audio (`loading_audio` in `config`, `loading_start` / `loading_stop` commands) on a dedicated `"loading-audio"` LiveKit track, independent of TTS

## [0.1.15] - 2026-04-30

### Features
Expand All @@ -17,8 +18,6 @@ All notable changes to Sayna will be documented in this file.

- Release v0.1.15 by @github-actions[bot]

- Release v0.1.15 by @github-actions[bot]


### CI/CD

Expand Down
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,7 +228,7 @@ See existing implementations for patterns:
- Receives: Config (optionally with `loading_audio`), audio data, speak commands, **clear command**, loading_start, loading_stop, send_message, sip_transfer
- Sends: Ready, STT results, TTS audio, messages, participant_connected, participant_disconnected, track_subscribed, tts_playback_complete, vad_event, error, sip_transfer_error
- **Clear command**: Immediately stops TTS and clears audio buffers (fire-and-forget, respects `allow_interruption` setting)
- **Loading commands**: `loading_start`/`loading_stop` control a looping loading-indicator audio clip on a dedicated `"loading-audio"` LiveKit track (fire-and-forget, independent of `speak`/`clear`)
- **Loading commands**: `loading_start`/`loading_stop` control a looping loading-indicator audio clip that is mixed into the single published audio track (fire-and-forget, independent of `speak`/`clear`). It is summed under TTS by the audio pump so any single-track subscriber (browser or SIP) hears it

### Webhooks
- `POST /livekit/webhook` - LiveKit event receiver (signature verified)
Expand Down
78 changes: 76 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

3 changes: 2 additions & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "sayna"
version = "0.1.16"
version = "0.1.15"
edition = "2024"
repository = "https://github.com/saynaai/sayna"
license = "Apache-2.0"
Expand Down Expand Up @@ -129,6 +129,7 @@ google-cloud-auth = "1.2.0"
tonic = { version = "0.11", features = ["tls-roots", "prost", "channel"] }
prost-types = "0.12"
async-stream = "0.3"
rubato = "3.0.0"

[dev-dependencies]
tower = { version = "0.5.2", features = ["util"] }
Expand Down
2 changes: 1 addition & 1 deletion Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,7 @@ RUN rm -rf /var/lib/apt/lists/* \
fi \
&& apt-get update -o Acquire::Retries=5 -o Acquire::http::No-Cache=true -o Acquire::Check-Valid-Until=false \
&& apt-get install -y --no-install-recommends -o Acquire::Retries=5 \
ca-certificates libssl3 libstdc++6 \
wget ca-certificates libssl3 libstdc++6 \
libva2 libva-drm2 \
libx11-6 libxext6 libxrandr2 libxcomposite1 libxdamage1 libxfixes3 \
libglib2.0-0 libgbm1 libdrm2 \
Expand Down
4 changes: 2 additions & 2 deletions docs/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -586,7 +586,7 @@ Configures audio processing and optional LiveKit mirroring. Must be the first me

**`loading_audio` configuration**

Decoded once at config time; the validated clip is held for the session and played on a dedicated `"loading-audio"` LiveKit track via the `loading_start` / `loading_stop` commands. Only 16-bit PCM is supported.
Decoded once at config time; the validated clip is resampled to the published track's format, held for the session, and mixed into the single published audio track via the `loading_start` / `loading_stop` commands. Only 16-bit PCM is supported.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
Expand Down Expand Up @@ -616,7 +616,7 @@ Immediately clears queued audio and LiveKit buffers. Useful for interruptions.
| `type` | string | `clear`. |

##### `loading_start`
Begins looping the configured loading indicator clip on a dedicated `"loading-audio"` LiveKit track. Requires `audio=true`, an active LiveKit room, and a `loading_audio` clip supplied in `config`. Fire-and-forget on success; idempotent if the loop is already running; failures emit an `error` message.
Begins looping the configured loading indicator clip, mixed into the single published audio track. Requires `audio=true`, an active LiveKit room, and a `loading_audio` clip supplied in `config`. Fire-and-forget on success; idempotent if the loop is already running; failures emit an `error` message.

| Field | Type | Description |
| --- | --- | --- |
Expand Down
15 changes: 6 additions & 9 deletions docs/livekit_integration.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ Sayna provides comprehensive LiveKit integration with automatic SIP infrastructu
## Table of Contents

1. [Architecture Overview](#architecture-overview)
2. [Published Audio Tracks](#published-audio-tracks)
2. [Published Audio Track](#published-audio-track)
3. [Configuration](#configuration)
4. [Inbound Webhooks (LiveKit → Sayna)](#inbound-webhooks-livekit--sayna)
5. [SIP Configuration & Auto-Provisioning](#sip-configuration--auto-provisioning)
Expand Down Expand Up @@ -92,20 +92,17 @@ Sayna provides comprehensive LiveKit integration with automatic SIP infrastructu

---

## Published Audio Tracks
## Published Audio Track

When a WebSocket session is audio-enabled and joins a LiveKit room, the Sayna agent participant publishes synthesized speech on an audio track named `"tts-audio"`.

If that session's `config` message also includes a `loading_audio` object (see [Loading Indicator](websocket.md#loading-indicator) in the WebSocket reference), the agent participant also publishes a **second** audio track named `"loading-audio"`, carrying the loading indicator sound. The agent then publishes **two** audio tracks:
When a WebSocket session is audio-enabled and joins a LiveKit room, the Sayna agent participant publishes **exactly one** audio track, named `"tts-audio"`. It carries the agent's full output — synthesized speech plus any loading-indicator sound mixed in.

| Track name | Carries |
|------------|---------|
| `"tts-audio"` | Synthesized speech (text-to-speech output). |
| `"loading-audio"` | The loading indicator clip, looped while the calling application is busy. |
| `"tts-audio"` | Synthesized speech, with the loading indicator clip mixed in while active. |

The two tracks are independent and can be audible at the same time. LiveKit client SDKs automatically play every subscribed audio track, so participants hear both without any extra client-side handling.
LiveKit is a Selective Forwarding Unit: it forwards each published track as an independent stream and never mixes audio server-side. Many subscribers play only a single audio track — custom browser clients that attach one `<audio>` element, and SIP bridges that downmix to a single phone stream — so publishing the loading indicator on a *second* track would not reliably reach them. Instead, Sayna mixes the loading clip into the one published track server-side (saturating sum, the same approach as LiveKit's own `AudioMixer`), so every single-track subscriber hears it. See [Loading Indicator](websocket.md#loading-indicator) in the WebSocket reference for the client-facing controls.

**Recordings:** Because `"loading-audio"` is a real published track, it is included in room-composite egress recordings alongside `"tts-audio"`. This is correct behavior — the recording faithfully reflects what the human participant heard.
**Recordings:** Because the loading indicator is part of the single `"tts-audio"` track, it is naturally included in room-composite egress recordings. The recording faithfully reflects what the human participant heard.

---

Expand Down
8 changes: 4 additions & 4 deletions docs/websocket.md
Original file line number Diff line number Diff line change
Expand Up @@ -435,7 +435,7 @@ If `allow_interruption: false` was set on the most recent speak command, the cle
```

**Behavior:**
- Starts the configured loading clip on a continuous, seamless loop on a dedicated `"loading-audio"` LiveKit track, separate from the `"tts-audio"` speech track.
- Starts the configured loading clip on a continuous, seamless loop, mixed into the single published audio track (`"tts-audio"`) and summed under any speech.
- If a loop is already running, the command is an idempotent no-op (the loop continues uninterrupted).
- The loading clip must have been supplied via the `loading_audio` object in the `config` message. See [Loading Indicator](#loading-indicator) for configuration and authoring guidance.

Expand Down Expand Up @@ -1513,7 +1513,7 @@ Use the token to connect to LiveKit room from web browsers or mobile apps using

The loading indicator is a short audio clip the server plays on a continuous, seamless loop into the LiveKit room while your application is busy ("thinking") and cannot yet answer. It is the audio equivalent of a spinner: it signals to the human participant that the agent is still working, instead of leaving the room silent.

The loading audio plays on its **own dedicated, second published LiveKit audio track** (`"loading-audio"`), separate from the speech track (`"tts-audio"`), so it never interferes with text-to-speech output.
The loading audio is **mixed into the single published audio track** (`"tts-audio"`), summed under text-to-speech output. LiveKit never mixes audio server-side and many subscribers (custom browser clients, SIP bridges) play only one audio track, so mixing into the one track is what lets every subscriber hear the indicator. See [LiveKit Integration](livekit_integration.md#published-audio-track) for details.

The feature has three parts:
1. **Configuration** — the `config` message gains an optional `loading_audio` object. The server decodes and validates the clip once and holds it for the session.
Expand Down Expand Up @@ -1595,9 +1595,9 @@ The server plays the clip **as supplied** (apart from the `volume` scaling and t

Loudness can be controlled in two ways: by authoring the clip quietly, and by setting the `volume` field (which scales the clip's sample amplitudes at config time). A LiveKit publisher cannot adjust a track's output volume after the fact, so these are the only loudness controls available.

### Separate Track Behavior
### Track and Mixing Behavior

The loading audio is published on its own LiveKit audio track (`"loading-audio"`), independent of the speech track (`"tts-audio"`). An audio-enabled session with `loading_audio` configured therefore publishes **two** audio tracks from the Sayna agent participant. See [LiveKit Integration](livekit_integration.md#published-audio-tracks) for details, including how the loading track appears in room-composite recordings.
The loading audio is mixed into the single published LiveKit audio track (`"tts-audio"`); an audio-enabled session publishes exactly one audio track from the Sayna agent participant whether or not `loading_audio` is configured. Mixing is a saturating sum of 16-bit samples (the same approach as LiveKit's own `AudioMixer`), so the indicator plays *under* speech rather than replacing it. See [LiveKit Integration](livekit_integration.md#published-audio-track) for why a single track is required and how the indicator appears in room-composite recordings.

After a LiveKit publisher-timeout reconnect, an in-progress loop stops (the old audio source is gone). The client must re-send `loading_start` to resume the loop.

Expand Down
15 changes: 13 additions & 2 deletions src/handlers/ws/config_handler.rs
Original file line number Diff line number Diff line change
Expand Up @@ -946,8 +946,19 @@ async fn initialize_livekit_client(
let mut livekit_client = LiveKitClient::new(livekit_config);

// Register the loading-indicator audio clip before connecting, if supplied.
if let Some(clip) = loading_clip {
livekit_client.set_loading_audio_clip(clip);
// Resampling to the published track format only fails on an internal/track
// misconfiguration, not bad client input: the decoder already enforces every
// precondition the resampler needs (mono/stereo, 8-48 kHz, whole frames).
// So this is treated as non-fatal — the indicator is simply unavailable, and
// the client is informed — rather than recorded as a `loading_audio`
// validation error (that channel is reserved for decode-time rejections).
if let Some(clip) = loading_clip
&& let Err(message) = livekit_client.set_loading_audio_clip(clip)
{
error!("Failed to prepare loading_audio: {}", message);
let _ = message_tx
.send(MessageRoute::Outgoing(OutgoingMessage::Error { message }))
.await;
}

// Set up audio callback to forward to STT processing
Expand Down
34 changes: 8 additions & 26 deletions src/handlers/ws/loading_handler.rs
Original file line number Diff line number Diff line change
Expand Up @@ -323,17 +323,13 @@ mod tests {
);
}

// This test constructs a real libwebrtc `NativeAudioSource` and so wears
// the `livekit_native_` quarantine prefix described in
// `src/livekit/client/tests.rs:499-506`: it is `#[ignore]`d out of the
// default `cargo test` run and executed isolated by the dedicated CI step
// in `.github/workflows/ci.yml`.
// `start_loading_audio` only needs an active connection and a configured
// clip (the pump consumes the overlay flag), so this success-path test needs
// no real audio source and runs in the default suite.
#[tokio::test]
#[ignore = "creates native libwebrtc objects; run isolated via the dedicated CI step (see ci.yml)"]
async fn livekit_native_handle_loading_start_message_success_is_silent() {
async fn test_handle_loading_start_message_success_is_silent() {
use crate::livekit::loading_clip::make_test_loading_clip;
use crate::livekit::{LiveKitClient, LiveKitConfig, sayna_audio_source_options};
use livekit::webrtc::audio_source::native::NativeAudioSource;
use crate::livekit::{LiveKitClient, LiveKitConfig};

let clip = make_test_loading_clip();
let mut client = LiveKitClient::new(LiveKitConfig {
Expand All @@ -348,15 +344,9 @@ mod tests {
listen_participants: vec![],
});
client.set_connected(true).await;
client.set_loading_audio_clip(clip);

let source = Arc::new(NativeAudioSource::new(
sayna_audio_source_options(),
16_000,
1,
100, // NativeAudioSource queue depth in ms (matches LOADING_AUDIO_QUEUE_SIZE_MS)
));
*client.loading_audio_source.lock().await = Some(source);
client
.set_loading_audio_clip(clip)
.expect("resampling the clip to the track format should succeed");

let mut state_inner = ConnectionState::new();
state_inner.set_audio_enabled(true);
Expand All @@ -373,14 +363,6 @@ mod tests {
message_rx.try_recv().is_err(),
"successful loading_start must not emit any WebSocket message"
);

// Tear down the spawned loop so the test does not leak a background task.
{
let guard = state.read().await;
if let Some(lk) = &guard.livekit_client {
lk.read().await.stop_loading_audio().await;
}
}
}

#[tokio::test]
Expand Down
Loading