CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Alkhouli, Tamer, Margatina, Katerina, Gung, James, Shu, Raphael, Zaghi, Claudia, Sunkara, Monica, Zhang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
by: K, Karthikeyan, et al.
Published: (2025)
by: K, Karthikeyan, et al.
Published: (2025)
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
by: Shu, Raphael, et al.
Published: (2024)
by: Shu, Raphael, et al.
Published: (2024)
Structured List-Grounded Question Answering
by: Sung, Mujeen, et al.
Published: (2024)
by: Sung, Mujeen, et al.
Published: (2024)
Controllable Conversational Theme Detection Track at DSTC 12
by: Shalyminov, Igor, et al.
Published: (2025)
by: Shalyminov, Igor, et al.
Published: (2025)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025)
by: Chai, Huacan, et al.
Published: (2025)
Eliciting Better Multilingual Structured Reasoning from LLMs through Code
by: Li, Bryan, et al.
Published: (2024)
by: Li, Bryan, et al.
Published: (2024)
Explicit Trait Inference for Multi-Agent Coordination
by: Abdurahman, Suhaib, et al.
Published: (2026)
by: Abdurahman, Suhaib, et al.
Published: (2026)
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
by: Vaccaro Jr, Michael, et al.
Published: (2024)
by: Vaccaro Jr, Michael, et al.
Published: (2024)
Personalized Turn-Level User Conversation Satisfaction Benchmark
by: Wang, Zhefan, et al.
Published: (2026)
by: Wang, Zhefan, et al.
Published: (2026)
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
by: Lumer, Elias, et al.
Published: (2025)
by: Lumer, Elias, et al.
Published: (2025)
RoundTable: Investigating Group Decision-Making Mechanism in Multi-Agent Collaboration
by: Cho, Young-Min, et al.
Published: (2024)
by: Cho, Young-Min, et al.
Published: (2024)
Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction
by: Myung, Jiyoon
Published: (2026)
by: Myung, Jiyoon
Published: (2026)
Applying General Turn-taking Models to Conversational Human-Robot Interaction
by: Skantze, Gabriel, et al.
Published: (2025)
by: Skantze, Gabriel, et al.
Published: (2025)
C-MTCSD: A Chinese Multi-Turn Conversational Stance Detection Dataset
by: Niu, Fuqiang, et al.
Published: (2025)
by: Niu, Fuqiang, et al.
Published: (2025)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
by: Juneja, Prerna, et al.
Published: (2026)
by: Juneja, Prerna, et al.
Published: (2026)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
by: Crouse, Maxwell, et al.
Published: (2026)
by: Crouse, Maxwell, et al.
Published: (2026)
DFlow: Diverse Dialogue Flow Simulation with Large Language Models
by: Du, Wanyu, et al.
Published: (2024)
by: Du, Wanyu, et al.
Published: (2024)
MemInsight: Autonomous Memory Augmentation for LLM Agents
by: Salama, Rana, et al.
Published: (2025)
by: Salama, Rana, et al.
Published: (2025)
Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious
by: Schessl, Ferdinand M.
Published: (2026)
by: Schessl, Ferdinand M.
Published: (2026)
Eliciting Behaviors in Multi-Turn Conversations
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
NaturalTurn: A Method to Segment Speech into Psychologically Meaningful Conversational Turns
by: Cooney, Gus, et al.
Published: (2024)
by: Cooney, Gus, et al.
Published: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
by: Horst, Nicola, et al.
Published: (2025)
by: Horst, Nicola, et al.
Published: (2025)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
by: Liu, Yu Lu, et al.
Published: (2026)
by: Liu, Yu Lu, et al.
Published: (2026)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
by: Badola, Kartikeya, et al.
Published: (2025)
by: Badola, Kartikeya, et al.
Published: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
by: Charpentier, Lucas, et al.
Published: (2025)
by: Charpentier, Lucas, et al.
Published: (2025)
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
by: Zhang, Michael J. Q., et al.
Published: (2024)
by: Zhang, Michael J. Q., et al.
Published: (2024)
LLMs Get Lost In Multi-Turn Conversation
by: Laban, Philippe, et al.
Published: (2025)
by: Laban, Philippe, et al.
Published: (2025)
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)
by: Rabinovich, Ella, et al.
Published: (2025)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Human Latency Conversational Turns for Spoken Avatar Systems
by: Jacoby, Derek, et al.
Published: (2024)
by: Jacoby, Derek, et al.
Published: (2024)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025)
by: Katsis, Yannis, et al.
Published: (2025)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
by: Choshen, Leshem, et al.
Published: (2026)
by: Choshen, Leshem, et al.
Published: (2026)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
by: Hu, Yuanzhe, et al.
Published: (2025)
by: Hu, Yuanzhe, et al.
Published: (2025)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
Similar Items
-
Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
by: Shen, Ming, et al.
Published: (2025) -
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
by: K, Karthikeyan, et al.
Published: (2025) -
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
by: Shu, Raphael, et al.
Published: (2024) -
Structured List-Grounded Question Answering
by: Sung, Mujeen, et al.
Published: (2024) -
Controllable Conversational Theme Detection Track at DSTC 12
by: Shalyminov, Igor, et al.
Published: (2025)