Arabic speech-to-text: four models against 40 hours of Gulf call audio
Word error rate per dialect on real contact centre recordings, including code-switched utterances, with the evaluation harness described in enough detail to reproduce.
Benchmarks, architecture decisions and operational write-ups. Mostly written by the engineers who shipped the thing being described.
Word error rate per dialect on real contact centre recordings, including code-switched utterances, with the evaluation harness described in enough detail to reproduce.
The latency arithmetic behind streaming recognition from the live call leg, and what breaks when you try it with a file-based pipeline.
How a single 100-point budget replaced three per-channel licences, and the queue behaviour we had to fix after it.
Practical notes on terminating trunks with regional carriers, codec negotiation and where jitter actually comes from.
The request lifecycle that resolves a tenant once, refuses disagreeing hints, and logs the attempt.
We share the Arabic evaluation setup with teams running their own comparison.