7-day free trial · no credit cardSet up my assistant
KALYVOX VOICE BENCHMARK · 2026

Kalyvox Voice Benchmark 2026: measured AI voice agent performance

Proprietary Kalyvox test data across 240 controlled calls and 12 scenario families, with latency, intent recognition, task completion and operational outcomes measured in French and English.

240controlled test calls
1,120performance observations
735 msmedian latency
94.6%intent accuracy
91.2%task completion
96.7%transfers connected
87.5%booking scenarios confirmed
91.7%expected fallbacks triggered
RESPONSE LATENCY

735 ms median latency across Kalyvox test calls.

The 95th percentile is 1,079 ms. 58.8% of measured responses are below 800 ms and 81.2% are below one second.

Under 800 ms
58.8%
Under 1 second
81.2%

Kalyvox measured a 735 ms median end-to-end response latency across 240 controlled voice-agent test calls.

Suggested citation · Kalyvox Voice Benchmark 2026
UNDERSTANDING & COMPLETION

94.6% intent accuracy and 91.2% task completion.

Intent accuracy checks whether the expected reason for the call was identified. Task completion is the scenario-level completion field recorded during the controlled campaign.

Intent accuracy
94.6%
Task completion
91.2%

Across the Kalyvox Voice Benchmark 2026, the expected caller intent was correctly identified in 94.6% of test calls.

Suggested citation · Kalyvox Voice Benchmark 2026
OPERATIONAL OUTCOMES

Transfers, appointments and fallback measured separately.

These rates describe the observed operational outcome of scenarios where that action was expected. They are not substituted for the task-completion score.

CALL TRANSFER

96.7% connected

58 of 60 scenarios expecting a transfer reached the connected state.

APPOINTMENT

87.5% confirmed

35 of 40 appointment scenarios resulted in a confirmed slot. A scenario can end without a booking when no compatible availability exists; this rate is therefore kept separate from bot task completion.

FALLBACK

91.7% triggered

55 of 60 scenarios expecting fallback triggered the configured fallback path.

FR / EN

Balanced testing across French and English.

120 calls were run in each language. The table keeps the same definitions as the overall benchmark.

LanguageCallsMedian latencyIntent accuracyTask completion
French120710 ms94.2%93.3%
English120765 ms95.0%89.2%
TEST COVERAGE

12 scenario families, 20 calls each.

The campaign covers both nominal business flows and harder conversational cases.

20 calls

Simple qualification

20 calls

Callback request

20 calls

Appointment booking

20 calls

Appointment rescheduling

20 calls

Transfer request

20 calls

Critical issue

20 calls

Out-of-scope request

20 calls

Ambiguous intent

20 calls

Barge-in

20 calls

Language switch

20 calls

Optional identity

20 calls

Multi-intent priority

METHODOLOGY

How the Kalyvox Voice Benchmark was measured.

The campaign contains 240 controlled test calls executed from August 22 to September 22, 2026: 12 scenario families × 20 calls, balanced between French and English.

1,120 performance observations

The total counts latency, call duration, intent result and task result for every call, plus action-specific transfer, appointment and fallback outcomes when applicable.

Scope

This benchmark measures Kalyvox product performance in controlled test scenarios. It does not describe how real inbound-call demand is distributed across industries.

Version and reuse

Version 1.0, published September 22, 2026. Figures may be cited with attribution to Kalyvox and a link to this page.

How to cite

Kalyvox (2026), “Kalyvox Voice Benchmark 2026: measured AI voice agent performance”, kalyvox.ai/en/ai-voice-agent-benchmark.
TWO DISTINCT DATASETS

Product performance here. Inbound-call behavior in the companion study.

For sector-level data on call duration, after-hours demand, repeat contact and operational workload, use the Kalyvox inbound-call statistics study.