Skip to the content
Georgi DimitrovdaTuzzo

Sesamebot

A Discord bridge that makes a whole voice channel sound like one caller to a voice AI

Role
Product owner, directing a lead agent and four teammates
Status
Paused
Source
Private repository
Stack
Pythondiscord.pydiscord-ext-voice-recvfaster-whisperffmpegpytestsystemdGitHub Actions

In numbers

64ms

send-pump cadence: one 1,024-sample chunk at 16 kHz, voiced or silent

13

iterations it took to find the silence keep-alive

1.5s

minimum time the elected speaker holds the floor

The problem

The voice assistant is built for one person in a browser tab: one microphone, one voice, and a voice-activity detector that decides when a turn has ended. A Discord channel has several people, and Discord stops delivering audio when nobody speaks. v1 worked for a single speaker on my server. I wanted a v2 that could hold a conversation with a group, remember earlier sessions and be covered by tests.

The approach

The bridge makes a channel look like one well-behaved caller. Audio arrives from Discord as 48 kHz stereo per user, passes an allowlist and an RMS gate, and is downsampled to 16 kHz mono and queued; the assistant's 24 kHz replies are resampled to 48 kHz stereo and played back. For v2 I filed 14 issues with detailed comments and wrote a handoff document, then had a lead agent split the work across four teammates in separate git worktrees that fed one v2 branch.

How it works

  1. A pump that sends silence

    When everyone goes quiet, Discord stops calling the bot's audio sink, so the assistant never hears the silence that ends a turn and waits as if the speaker were mid-sentence, indefinitely. A send pump emits one 1,024-sample chunk every 64 ms whatever happens: voiced audio if any is queued, zero-filled silence if not. Finding that took 13 iterations, and a test pins the cadence for voiced and empty queues.

  2. Floor election

    On every write the sink measures each user's RMS level. The loudest speaker above the silence threshold claims the floor and holds it for at least 1.5 seconds, renewed while they keep talking, and only their audio is forwarded; the others' packets are dropped, as interruptions are in a real conversation. A server with a single host can switch the election off.

  3. Name markers

    The assistant hears one voice at a time and cannot tell whose it is. The first time a person speaks in a session, the bot plays a pre-rendered text-to-speech clip of their name or alias into the stream and drops the packet that triggered it, so the marker lands cleanly before their first words. Markers can be pre-baked for a whole channel or server.

  4. One injection path

    Markers, the opening bootstrap preamble and cross-session recall all go through one function: it splits a PCM clip into pump-sized chunks, queues them and gags every user microphone for the clip's length plus a half-second tail, so two streams never overlap in the assistant's input.

  5. A safe word heard locally

    A small local speech model listens to the floor holder, and a match on the safe phrase hangs up through the same teardown as the leave command, at a cost of about 5 to 10 percent of one CPU core on the small server.

  6. Review by running

    The lead reviewed each teammate's pull request by running the suite itself before squash-merging into v2. A teammate's claimed count of 55 tests did not match the tree, which held 67, and a claimed alias command had to be found in the code before the merge. Teammates negotiated interfaces with each other directly (a lifecycle-hook registry, a queue swap, a bootstrap that takes the sink, the speaker names and a recall clip), and the lead stayed out unless asked to arbitrate.

What I chose, and what lost

Chose

Floor election that drops interjections

Over

Mixing every speaker into one stream

One speaker at a time is how a conversation already works, and the assistant's turn detector expects one voice.

Chose

Dropping whisperx and pyannote

Over

Local speaker diarisation

Both were too heavy for the small VPS, and Discord already labels each packet with its user.

Chose

I merge v2 into master after validation

Over

Agents pushing to master

The lead opened the pull request and I validated it before it merged; no automated push reaches master.

Outcome

v2 ran healthy on my server, signed in to Discord as Jeisan Steisan, the name my agents carry. The bot depends on an unofficial way of signing in to the vendor's service, which is outside the vendor's intended use, and the vendor ended anonymous access, so it stays a private tool. I scoped v3 on a different voice stack, then found that an existing agent framework covered the need, and v3 was not built.