Munsit is an Arabic Voice AI platform built in the UAE, offering Arabic speech-to-text via REST and WebSocket streaming behind a single API key, documented with a published OpenAPI specification and a WebSocket protocol specification.
API surface:
•Transcribe: POST /audio/transcribe send a file, receive a transcript
•Stream in real time: WebSocket at /websocket/speech-to-text transcripts arrive as the speaker talks
•Speaker diarization: POST /audio/diarization/transcribe
•Diarization + sentiment: per-speaker sentiment analysis layered on a diarized transcript
•Minutes of meetings: POST /minutes-of-meeting/transcribe send a recording, receive structured minutes directly
•Keyword extraction and translation on transcripts and meeting outputs
•Voice isolation: POST /denoise clean noisy contact-center or field audio before transcription, in the speech to text API
Dialect handling is automatic, with no locale parameter to configure. The model covers 25+ dialects Gulf varieties (Emirati, Khaleeji, Najdi, Hijazi), Levantine, Egyptian, Sudanese, Iraqi, North African dialects, and MSA and identifies the spoken variety on its own, including handling code-switching between Arabic and English within the same conversation. This matters specifically for the locale-parameter problem described above: a single call that shifts between dialects and English doesn’t require the caller to guess or declare a variety in advance.
The understanding layer sits behind the same key. Diarization, per-speaker sentiment, meeting minutes, keyword extraction, translation, and voice isolation are all reachable with the same API key used for transcription for teams building call analytics or compliance review, this collapses what is normally a three-vendor pipeline (ASR, NLP, audio cleanup) into one.
Real-time streaming is available over a dedicated WebSocket (/websocket/speech-to-text), producing partial transcripts as the speaker talks, suitable for live call monitoring, voice agents, and real-time dashboards. Batch file processing runs through the REST endpoint for recorded audio.
Voice agent framework integrations: drop-in plugins for LiveKit, Pipecat, VAPI, and Ultravox, so teams building Arabic voice agents on standard frameworks integrate without writing custom WebSocket handling.
Deployment options: cloud API, sovereign cloud (VPC), on-premises for air-gapped government and banking environments, and on-device processing.
Authentication: a single API key (x-api-key header) covers transcription, the understanding layer, and Munsit’s text-to-speech product.
Pricing: free credits on signup, no card required. Paid plans from $8/month (billed annually) with 200,000 credits/month; enterprise pricing for sovereign deployment and high volume. (Verify current rates and the credits-to-audio-duration conversion directly at munsit.com/pricing, since usage-based pricing details change.)