وكلاء صوتيون بالعربية
Arabic voice agents that answer under 800 milliseconds
Deepgram streaming recognition, GPT-4o reasoning and ElevenLabs synthesis, wired so the caller hears a response before a human would have finished reading the intent. Each tier has an in-boundary counterpart.
Where the 800 milliseconds go
Move the sliders to see how recognition, reasoning and synthesis compose into a single perceived turn, against a legacy IVR baseline.
Response time against unit cost
Over 8 conversational turns that is 13.8s of dead air removed from every call.
Baseline 0.95 USD per agent minute · AI pipeline 0.0095 USD per minute · currency pegs applied at par.
Interactive model. Production p95 measured from end of caller speech to first synthesised audio frame.
Three tiers, each swappable
Recognition
Deepgram streaming transcription with dialect models and per-speaker tagging on forked legs. Self-hosted faster-whisper for tenants that cannot egress audio.
Reasoning
GPT-4o-mini for turn-level intent, tool calls and guardrails against tenant knowledge. vLLM-served open weights for the sovereign tier.
Synthesis
ElevenLabs multilingual voices with Arabic prosody control, or Azure Neural and in-boundary TTS where residency requires it.
Why dialect is the whole product
A caller in Sharjah does not speak Modern Standard Arabic, and a model trained on MSA news audio will mishear them. Recognition is evaluated per dialect against tenant call recordings before a flow goes live.
Code-switching is the normal case, not an edge case. A single sentence carries an Arabic verb, an English product name and a number read in either language; the pipeline is tuned to keep all three.
- Gulf, Levantine, Egyptian and MSA recognition profiles
- Arabic–English code-switching inside one utterance
- Barge-in: the caller can interrupt and the turn re-plans
- Numbers, dates and Emirates ID formats normalised before tool calls
Hear it answer in your dialect
We run a live call on your own recordings during the demo, with recognition accuracy reported per dialect.