A LiveKit voice agent follows an STT → LLM → TTS pipeline inside an AgentSession: incoming audio is transcribed in real time, the transcript goes to a language model for reasoning, and the model’s reply is synthesized back to speech and streamed to the caller with LiveKit’s turn-detection logic deciding when the caller has actually finished speaking versus just pausing mid-thought. This guide uses Munsit for both the STT and TTS legs and OpenAI for the LLM leg, which is the configuration documented in Munsit’s own LiveKit integration guide.
Prerequisites
• Python 3.10 or later
• A Munsit account and an API key (generated from Munsit’s API Keys dashboard)
• A LiveKit Cloud account, or a self-hosted LiveKit server
• An LLM provider key (this guide uses OpenAI, but any LLM plugin LiveKit Agents supports will work)
Step 1: Install and Authenticate
Install the plugin from PyPI, or from source if you want the latest unreleased changes:
# From PyPI
pip install livekit-plugins-munsit
# From source
git clone https://github.com/CNTXTFZCO0/livekit-plugins-munsit.git
cd livekit-plugins-munsit
pip install -e .
Both munsit.STT and munsit.TTS read MUNSIT_API_KEY automatically from the environment. Put your credentials in a .env.local file alongside your LiveKit and LLM keys:
# Munsit speech-to-text and text-to-speech
MUNSIT_API_KEY=your_MUNSIT_API_KEY_here
# LiveKit Configuration
LIVEKIT_URL=wss://your-livekit-server.com
LIVEKIT_API_KEY=your_livekit_api_key
LIVEKIT_API_SECRET=your_livekit_api_secret
# LLM Configuration (for the agent's brain)
OPENAI_API_KEY=your_openai_api_key
Your Munsit API key is shown once at creation. Keep .env.local out of version control, use environment variables (not hardcoded strings) in production, and rotate keys periodically standard key hygiene that matters more here because this key authenticates both the STT and TTS legs of a live, customer-facing pipeline.
Step 2: Write the Agent
Create arabic_agent.py:
"""
Arabic voice assistant Munsit STT + Munsit TTS
"""
from dotenv import load_dotenv
from livekit import agents
from livekit.agents import Agent, AgentServer, AgentSession
from livekit.plugins import munsit, openai, silero
load_dotenv(".env.local")
class ArabicAssistant(Agent):
"""Arabic-speaking voice assistant"""
def __init__(self) -> None:
super().__init__(
instructions="""
أنت مساعد صوتي ذكي يتحدث العربية بطلاقة.
مهمتك مساعدة المستخدمين بالإجابة على أسئلتهم بطريقة واضحة ومفيدة.
كن ودوداً ومحترماً في تعاملك.
"""
)
server = AgentServer()
@server.rtc_session()
async def my_agent(ctx: agents.JobContext):
session = AgentSession(
stt=munsit.STT(model="munsit-en-ar"), # Arabic + English code-switch
llm=openai.LLM(model="gpt-4o", temperature=0.7),
tts=munsit.TTS(
voice_id="ar-uae-male-1", # Copy IDs from the Voice Library
model="faseeh-v1-preview",
stability=0.75,
speed=1.0,
),
vad=silero.VAD.load(activation_threshold=0.6, min_speech_duration=0.3),
)
await session.start(room=ctx.room, agent=ArabicAssistant())
await session.generate_reply(
instructions="رحب بالمستخدم باللغة العربية وقدم نفسك كمساعد ذكي جاهز للمساعدة."
)
if __name__ == "__main__":
agents.cli.run_app(server)
The Silero VAD thresholds here (activation_threshold=0.6, min_speech_duration=0.3) are tuned for real-world microphones, where post-echo-cancellation audio leaking back from the agent’s own speaker can otherwise be misread as the caller speaking the “microphone feedback loop” failure mode covered in troubleshooting below.
Step 3: Run It
# Dev mode: starts a local LiveKit server, launches the agent, gives you a test URL
python arabic_agent.py dev
# Production: deploy against your LiveKit Cloud or self-hosted server
python arabic_agent.py start
In dev mode, open the test URL in a browser, allow microphone access, and speak Arabic the agent should respond in the configured Munsit voice. For the fastest debugging loop, use console mode, which runs locally against your microphone and prints transcripts to stdout without needing a LiveKit server at all.
Step 4: Configure Speech-to-Text
munsit.STT() defaults to streaming mode: it holds a live WebSocket to Munsit, emits interim transcripts while the caller is still talking, and finalizes each turn with word-level timestamps the moment server-side turn detection fires the right mode for a live voice agent. Pass mode="batch" to instead buffer a full utterance and POST it as a single request, which suits pipelines transcribing recorded audio rather than live conversation.
| Mode |
Best For |
Endpoint |
Typical Latency |
|
streaming (default)
|
Voice agents, live captions
|
WS /api/v1/listen
|
~0.7s median to first partial; final ~1.2s after speech ends
|
|
batch
|
Recorded audio, one-request-per-utterance pipelines
|
POST /api/v1/audio/transcribe
|
~1–2s for short utterances (VAD + upload + processing)
|
from livekit.plugins import munsit
streaming_stt = munsit.STT()
batch_stt = munsit.STT(mode="batch")
Note: STT.recognize(audio_buffer) always hits the batch HTTP endpoint directly, regardless of what mode the instance was configured with use it when you have a recorded buffer (a voicemail, an uploaded file) rather than a live stream.
Choosing a model. Pin the model explicitly rather than relying on the plugin default, so a future default change can’t silently alter your agent’s behavior:
| Model |
Use Case |
|
munsit-en-ar (plugin default)
|
Mixed Arabic-English speech with code-switching
|
|
munsit
|
Pure Arabic recognition; fastest, and the only model that supports custom vocabulary (hotwords)
|
A few constructor parameters worth knowing beyond the defaults: hotwords accepts up to 200 custom vocabulary entries (rare terms or names, ≤40 characters each) but only works on the munsit model passing it with munsit-en-ar is silently ignored with a warning. sample_rate defaults to 16000 and is just an input hint; the streaming endpoint negotiates 8000/16000 with the server and resamples anything else automatically, which matters because LiveKit tracks are typically 48000. auth_method accepts header (default, x-api-key), bearer, or query the query option exists specifically for deployments behind a proxy that strips custom headers.
Step 5: Configure Text-to-Speech
TTS streaming is always on and the agent starts speaking as soon as the first audio chunk is ready, rather than waiting for the full reply to render. Pick a voice from Munsit’s Voice Library by listening to samples and copying the voice ID.
| Parameter |
Values |
What it does |
| voice_id |
e.g. ar-uae-male-1 |
The Arabic voice to synthesize with |
| model |
faseeh-v1-preview |
Munsit’s Arabic voice synthesis model |
| stability |
0.0–1.0 |
Voice consistency: 0.0–0.4 more expressive but can hallucinate; 0.5–0.7 balanced (recommended); 0.8–1.0 very consistent, less variation
|
| speed |
0.7–1.2 |
Speech rate: 0.7–0.9 slower/clearer for complex content; 1.0 normal; 1.1–1.2 faster
|
| sample_rate |
8000–48000 |
Output PCM rate; defaults to 48000, the engine’s native rate, so WebRTC audio needs no resampling. Lower it only for narrowband transport like telephony
|
| dialect |
e.g. emirati |
Optional pronunciation hint for the selected voice
|
Rule of thumb for stability: chatbots 0.6–0.8, professional/formal applications 0.8–1.0, creative content 0.3–0.5. Voice settings can also change at runtime useful for switching tone or dialect mid-conversation based on context:
# Start with one voice
tts = munsit.TTS(voice_id="ar-uae-male-1", speed=1.0)
# Switch to a different voice based on context
tts.update_options(
voice_id="ar-hijazi-female-2",
stability=0.8,
speed=1.0,
)
Step 6: Tune Endpointing
In streaming mode, the server decides when a turn ends: it waits for endpointing_ms of silence (default 800ms, tunable 100–5000ms), and with smart_turn enabled (default) a semantic turn-completion model must also agree the speaker has actually finished, so a caller pausing mid-thought isn’t cut off mid-sentence. A turn always ends after twice the silence window regardless, as a hard ceiling.
# Snappier finals: shorter silence window, semantic gating still on
stt = munsit.STT(
mode="streaming",
endpointing_ms=300,
smart_turn=True,
)
# Retune mid-session on a live stream (100–5000ms)
stream = stt.stream()
await stream.configure_endpointing(800)
For a recorded buffer rather than a live stream, use recognize() directly:
from livekit import rtc
from livekit.plugins import munsit
stt = munsit.STT()
frames = [...] # list of rtc.AudioFrame
combined = rtc.combine_audio_frames(frames)
result = await stt.recognize(combined)
print(result.alternatives[0].text)
for word in result.alternatives[0].words:
print(f"{word.start_time:.2f}s -> {word.end_time:.2f}s {word}")
Step 7: Track Turn Metrics
Each conversation turn carries timing data on its ChatMessage. Subscribing to conversation_item_added gives transcription delay, end-of-turn delay, and downstream LLM/TTS timing useful for diagnosing where latency is actually coming from in production rather than guessing:
from livekit.agents import ChatMessage
@session.on("conversation_item_added")
def on_item(event):
msg = event.item
if not isinstance(msg, ChatMessage):
return
metrics = msg.metrics or {}
if msg.role == "user":
transcription_delay = metrics.get("transcription_delay")
end_of_turn_delay = metrics.get("end_of_turn_delay")
if transcription_delay is not None:
print(f"STT delay: {transcription_delay * 1000:.0f} ms")
if end_of_turn_delay is not None:
print(f"EOU delay: {end_of_turn_delay * 1000:.0f} ms")
elif msg.role == "assistant":
llm_ttft = metrics.get("llm_node_ttft")
tts_ttfb = metrics.get("tts_node_ttfb")
if llm_ttft:
print(f"LLM TTFT: {llm_ttft * 1000:.0f} ms")
if tts_ttfb:
print(f"TTS TTFB: {tts_ttfb * 1000:.0f} ms")
The older metrics_collected event is deprecated in favor of this approach.