Introducing ฯ-Entity, a benchmark for how voice agents collect, verify, and correct the details callers give them.
In ฯ-Voice, capturing names and other identifying information was a recurring bottleneck: get a detail wrong, and authentication fails before the agent can help. ฯ-Entity breaks that exchange down. Can an agent collect one exact value, recognize when it needs to check what it heard, and repair a mistake?
We test 200 tasks across ten entity types, including names, phone numbers, addresses, and emails. Each call asks for a single field, with a cooperative caller who spells accurately when asked. Most calls last less than two minutes. And even here, reliability is far from solved.
When agents choose their own verification strategy, only 14โ41% of tasks succeed across all tested conditions. A brute-force approachโprescribing spelling, read-back, and confirmationโraises that to 37โ54%, but makes calls longer. The challenge is finding the balance: knowing when to trust a capture and when to spend more effort checking it.
Across systems, agents recognize hard and unfamiliar entities and verify them more carefully. But they do not increase verification for the caller voices they struggle with most, and their response to noise is inconsistent. And verification often fails to produce a repair: just 24โ37% of wrong captures that agents verify end up correct. Agents need to get better at both deciding when to check and using the answer to fix the record.
Weโre introducing ฯ-Entity alongside ฯ-i18n, and plan to extend these entity-collection tasks across languages with the upcoming codebase release. Together, they help us measure whether voice agents can both handle the details accurately and deliver a conversation that feels native.
Thanks to Victor Barres at
@mercor for collaborating on this.
Papers: ฯ-Entity
arxiv.org/abs/2609.13602v1, ฯ-i18n
arxiv.org/abs/2609.35820v1
Codebase and multilingual extension coming soon!