VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Dhruv, Shukla, Harshit, Rajeev, Gautam, Kulkarni, Ashish, Khatri, Chandra, Agarwal, Shubham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
by: Jain, Purvam, et al.
Published: (2026)
by: Jain, Purvam, et al.
Published: (2026)
VoiceBench: Benchmarking LLM-Based Voice Assistants
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
by: Manoj, Guduru, et al.
Published: (2025)
by: Manoj, Guduru, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
by: Gautam, Dhruv, et al.
Published: (2025)
by: Gautam, Dhruv, et al.
Published: (2025)
Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
by: Das, Devleena, et al.
Published: (2025)
by: Das, Devleena, et al.
Published: (2025)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
by: Patel, Shubham, et al.
Published: (2024)
by: Patel, Shubham, et al.
Published: (2024)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
by: Atreja, Dhruv
Published: (2025)
by: Atreja, Dhruv
Published: (2025)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Voices of Her: Analyzing Gender Differences in the AI Publication World
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
by: Nix, Arne, et al.
Published: (2026)
by: Nix, Arne, et al.
Published: (2026)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
by: Thonet, Thibaut, et al.
Published: (2024)
by: Thonet, Thibaut, et al.
Published: (2024)
Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies
by: Kulkarni, Siddhant, et al.
Published: (2026)
by: Kulkarni, Siddhant, et al.
Published: (2026)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
by: Guo, Yiju, et al.
Published: (2026)
by: Guo, Yiju, et al.
Published: (2026)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
by: Mahapatra, Debabrata, et al.
Published: (2025)
by: Mahapatra, Debabrata, et al.
Published: (2025)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
by: Agarwal, Dhruv, et al.
Published: (2023)
by: Agarwal, Dhruv, et al.
Published: (2023)
Krutrim LLM: Multilingual Foundational Model for over a Billion People
by: Kallappa, Aditya, et al.
Published: (2025)
by: Kallappa, Aditya, et al.
Published: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Proto Successor Measure: Representing the Behavior Space of an RL Agent
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
by: Mazumder, Aritra, et al.
Published: (2026)
by: Mazumder, Aritra, et al.
Published: (2026)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants
by: Herrera, Alejandro Breen, et al.
Published: (2026)
by: Herrera, Alejandro Breen, et al.
Published: (2026)
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
by: Phan, Peter, et al.
Published: (2025)
by: Phan, Peter, et al.
Published: (2025)
JAF: Judge Agent Forest
by: Garg, Sahil, et al.
Published: (2026)
by: Garg, Sahil, et al.
Published: (2026)
DataSciBench: An LLM Agent Benchmark for Data Science
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
TaeBench: Improving Quality of Toxic Adversarial Examples
by: Zhu, Xuan, et al.
Published: (2024)
by: Zhu, Xuan, et al.
Published: (2024)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
Similar Items
-
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025) -
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024) -
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024) -
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
by: Bogavelli, Tara, et al.
Published: (2026) -
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
by: Jain, Purvam, et al.
Published: (2026)