Command Palette
Search for a command to run...
EVA-Bench end-to-end Speech Agent Benchmark Dataset
EVA-Bench is an end-to-end speech agent evaluation benchmark dataset released by ServiceNow. Related research papers include... EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsIt aims to evaluate the task completion ability and interactive experience of voice intelligent agents, and is widely used in the simulation dialogue generation and quality measurement of conversational voice intelligent agents. This dataset contains 213 scenarios, covering three enterprise domains: aviation customer service management, medical human resources services, and enterprise IT service management. It involves 39 types of workflows and 121 tool calls, and supports a comprehensive evaluation of multi-turn voice interaction, tool calls, task completion status, and robustness of voice scenarios.
Enterprise Domain Classification
- Airline Customer Service Management (CSM): Covers 50 scenarios, including abnormal flight rebooking, voluntary changes, cancellations, same-day standby, compensation voucher issuance, and adversarial users.
- Healthcare Human Resources Services (HRSD): Covering 83 scenarios, including onboarding and qualification verification of medical personnel, license and DEA registration verification, OTP secondary authentication, leave and special applications, dual/triple intent requests, and adversarial users, etc.
- Enterprise IT Service Management (ITSM): Covers 80 scenarios, including incident classification and handling, escalation control, change and issue management, asset and access control, multi-level authentication, single to four-intention compound requests, and adversarial users.

Flowchart
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.