Cekura Bench: Speech-to-speech model benchmarks on live phone calls
The 82 scenarios split into 59 appointment-booking cases and 23 Medicare cases. Each realtime model gets the same prompt, tools and connection inside a shared open-source Pipecat pipeline. The benchmark site now lists 11 models and a cascade baseline across 2,952 calls. The configurations tested include GPT Realtime 2.1 and 2.1 Mini, Gemini 3.1 Flash Live, Grok Voice Think Fast 2.0, Nova 2 Sonic, Qwen, and Phonic v1. The harness code is public on GitHub. An earlier Cekura board used the same suite to compare full voice-agent platforms. In that ranking, Retell led, passing 62 of 82 scenarios on all three runs, and Gemini Live came last at 30.49%. Rankings are based on repeatable pass^3 reliability, so a scenario counts as passed only if it succeeds every time. Most other speech-to-speech evaluations, such as VoiceBench or the Artificial Analysis speech-to-speech leaderboard, test models outside a live phone call. For founders, the useful signal is the gap between completing a call once and doing it reliably. Telnyx completed tasks on 97.56% of calls but passed only 68.29% of scenarios on all three runs. Results from a vendor's own benchmark deserve a close look, but the public transcripts let teams check the claims before choosing a model.