Introducing Sales Agent Eval by
@voicearena_ai
We're inviting the community to share feedback on the methodology for a real-world benchmark for sales conversation agents. To our knowledge, this is the first live benchmark that ranks voice agents by revenue from real sales against a human baseline.
Overview
We evaluate voice sales agents in a live experiment with real customers and real money. A company set up under Voice Arena sells an online course in the United States. Leads generated through Meta ads are randomly split between K commercial AI voice agents and three human salespeople. Participants are ranked by the revenue they generate, and the leaderboard is updated live and finalized at the end of each day.
Product
Participants sell a self-paced online course on AI tools for work, including ChatGPT, Excel and productivity tools. We chose this category because it has broad consumer appeal, strong current demand and relatively low production costs.
Online courses can be bought in a single call at an impulse price, are delivered digitally with no logistics, and require genuine objection handling. Course telesales is also an established industry, so the human baseline reflects real practice. The course is sold under identical offer rules for every participant (list price, a fixed discount ladder and a refund policy), which stay constant for the whole season.
Participants
There are two kinds of participants: K AI agents and three human salespeople. The AI agents are the systems under evaluation, and the humans provide the baseline.
The AI participants are commercial vendor stacks. To be included, a vendor must support outbound calling in US English, allow tool calls during the conversation to send payment links, expose per-call cost and recordings, and allow its configuration to be pinned for the season. We configure every agent from one shared sales brief, knowledge base, tool set and set of compliance rules. Only platform-specific settings such as voice, model and latency parameters are tuned, and each agent receives an equal tuning budget of pilot calls on a separate pool of leads. Configurations are then frozen and snapshotted (models, voices, prompt hash and platform version), and any change creates a new entry that starts from zero leads.
The human participants are three salespeople with prior telesales experience. They receive the same materials and offer rules as the AI agents, are paid a salary plus commission, and do not see the leaderboard during the season. Each salesperson is reported individually, and the three are also reported together as a pooled human result.
Assignment
All leads received each day are randomly assigned across the AI agents and the three human salespeople. Each participant receives the same number of new leads per day, capped by human calling capacity at approximately 25 new leads per participant. Because everyone draws from the same daily pool, every participant gets leads of the same quality, and version 1 does not count the advertising cost of acquiring them.
Calling in the United States
Every lead gives prior express written consent on the lead form to be called about the course, including by an AI voice. The Federal Communications Commission confirmed in 2024 that AI-generated voices count as an "artificial voice" under the Telephone Consumer Protection Act, so AI calls require the called party's prior consent. The same consent is collected from every lead, so AI and human participants call from an identical pool.
All participants call only between 8 a.m. and 9 p.m. in the lead's local time, or within a narrower window where state law requires one. Every call opens with the brand name and a recording notice, because several states, including California, Florida and Pennsylvania, require every party to consent to a recording. AI participants also disclose that the caller is an AI. Payment links and follow-ups are sent by SMS from a shared registered number, and an opt-out on any channel stops contact from every participant.
Cost accounting
Version 1 ranks on total revenue, but we also measure what each participant costs to run, so revenue can be read alongside it. For AI participants, the operating cost is everything billed for running the agent: speech recognition, language model, speech synthesis, platform fees, telephony minutes including failed attempts and retries, and SMS. Vendor usage is priced at public list rates snapshotted at the start of the season, and any mid-season price change is noted but not applied. For human participants, the operating cost is each salesperson's salary and commission over the season. Operating cost per lead is total operating cost divided by the number of leads assigned. Advertising and one-time setup and tuning costs are not included.
Ranking and leaderboard
Participants are ranked by total revenue, which compares fairly because everyone gets the same number of leads of the same quality each day. Revenue is the amount paid on each lead's first purchase, excluding sales tax. Refunds and chargebacks are not deducted in version 1, so a sale counts at the moment it is made, regardless of what happens afterwards. Operating cost and revenue after operating cost are shown beside total revenue but don't affect the rank.
A purchase counts if it is made within 7 days of the lead being assigned, so a lead's outcome is final after 7 days. The board is finalized at the end of each day, and stays provisional until every participant has 500 leads past the 7-day mark.
Each season runs for 6 weeks of lead assignment, followed by a 7-day settlement period, after which the official ranking is computed.
Get involved
We want this methodology to be shaped by the people building and buying sales agents. Visit Voice Arena at the link below and fill in the request form to:
- Help shape the methodology
- Read the full detailed version
- Submit your agent for evaluation