Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Zhaoyi Joey, Shourya, Tanya, Wang, Yingfan, Roy, Shamik, Kumar, Vinayshekhar Bannihatti, Gangadharaiah, Rashmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Prompt KV Cache: Where It Becomes Dispensable
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026)
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026)
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
by: Zhou, Hanhan, et al.
Published: (2026)
by: Zhou, Hanhan, et al.
Published: (2026)
FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis
by: Makhija, Disha, et al.
Published: (2025)
by: Makhija, Disha, et al.
Published: (2025)
Constrained Decoding with Speculative Lookaheads
by: Nakshatri, Nishanth, et al.
Published: (2024)
by: Nakshatri, Nishanth, et al.
Published: (2024)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
by: Shim, Jeonghoon, et al.
Published: (2025)
by: Shim, Jeonghoon, et al.
Published: (2025)
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
by: Agarwal, Dhruv, et al.
Published: (2023)
by: Agarwal, Dhruv, et al.
Published: (2023)
Multi-Faceted Evaluation of Modeling Languages for Augmented Reality Applications -- The Case of ARWFML
by: Muff, Fabian, et al.
Published: (2024)
by: Muff, Fabian, et al.
Published: (2024)
Multi-Facet Counterfactual Learning for Content Quality Evaluation
by: Zheng, Jiasheng, et al.
Published: (2024)
by: Zheng, Jiasheng, et al.
Published: (2024)
Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework
by: Xu, Kaishuai, et al.
Published: (2025)
by: Xu, Kaishuai, et al.
Published: (2025)
Unlocking Efficiency: Adaptive Masking for Gene Transformer Models
by: Roy, Soumyadeep, et al.
Published: (2024)
by: Roy, Soumyadeep, et al.
Published: (2024)
Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
by: Kumar, Vanya Bannihatti, et al.
Published: (2025)
by: Kumar, Vanya Bannihatti, et al.
Published: (2025)
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale
by: Cho, Jungho, et al.
Published: (2026)
by: Cho, Jungho, et al.
Published: (2026)
Loops On Retrieval Augmented Generation (LoRAG)
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
by: Khandelwal, Dinesh, et al.
Published: (2026)
by: Khandelwal, Dinesh, et al.
Published: (2026)
Model Fusion with Multi-LoRA Inference for Tool-Enhanced Game Dialogue Agents
by: Wang, Kangxu, et al.
Published: (2025)
by: Wang, Kangxu, et al.
Published: (2025)
PeeriScope: A Multi-Faceted Framework for Evaluating Peer Review Quality
by: Ebrahimi, Sajad, et al.
Published: (2026)
by: Ebrahimi, Sajad, et al.
Published: (2026)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
by: Jang, Kyochul, et al.
Published: (2025)
by: Jang, Kyochul, et al.
Published: (2025)
Language Models Entangle Language and Culture
by: Jain, Shourya, et al.
Published: (2026)
by: Jain, Shourya, et al.
Published: (2026)
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
by: Jia, Qi, et al.
Published: (2026)
by: Jia, Qi, et al.
Published: (2026)
MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation
by: He, Junqing, et al.
Published: (2024)
by: He, Junqing, et al.
Published: (2024)
A Unified Data Augmentation Framework for Low-Resource Multi-Domain Dialogue Generation
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
by: Du, Yaxin, et al.
Published: (2025)
by: Du, Yaxin, et al.
Published: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
by: Kuo, Tzu-Lin, et al.
Published: (2024)
by: Kuo, Tzu-Lin, et al.
Published: (2024)
Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Multi-Facet Blending for Faceted Query-by-Example Retrieval
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Data Augmentation of Multi-turn Psychological Dialogue via Knowledge-driven Progressive Thought Prompting
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
by: Park, Minhu, et al.
Published: (2025)
by: Park, Minhu, et al.
Published: (2025)
Efficient Tool-Calling Multi-Expert NPC Agent for Commonsense Persona-Grounded Dialogue
by: Nuriyev, Mahammad
Published: (2025)
by: Nuriyev, Mahammad
Published: (2025)
Tool-to-Agent Retrieval: Bridging Tools and Agents for Scalable LLM Multi-Agent Systems
by: Lumer, Elias, et al.
Published: (2025)
by: Lumer, Elias, et al.
Published: (2025)
Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
by: Vasselli, Justin, et al.
Published: (2025)
by: Vasselli, Justin, et al.
Published: (2025)
FLAP: Flow-Adhering Planning with Constrained Decoding in LLMs
by: Roy, Shamik, et al.
Published: (2024)
by: Roy, Shamik, et al.
Published: (2024)
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
by: Brehme, Lorenz, et al.
Published: (2026)
by: Brehme, Lorenz, et al.
Published: (2026)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
by: Peng, Bo, et al.
Published: (2025)
by: Peng, Bo, et al.
Published: (2025)
Similar Items
-
Probing the Prompt KV Cache: Where It Becomes Dispensable
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026) -
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025) -
Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language
by: Kumar, Vinayshekhar Bannihatti, et al.
Published: (2026) -
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
by: Zhou, Hanhan, et al.
Published: (2026) -
FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance
by: Kang, Mintong, et al.
Published: (2025)