Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Choubey, Prafulla Kumar, Peng, Xiangyu, Bhagavath, Shilpa, Xiong, Caiming, Pentyala, Shiva Kumar, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking Deep Search over Heterogeneous Enterprise Data
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024)
by: Xie, Kaige, et al.
Published: (2024)
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
by: Zhang, Nan, et al.
Published: (2024)
by: Zhang, Nan, et al.
Published: (2024)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
Nudging the Boundaries of LLM Reasoning
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents
by: Zhou, Yuqing, et al.
Published: (2026)
by: Zhou, Yuqing, et al.
Published: (2026)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
by: Saraogi, Devesh, et al.
Published: (2025)
by: Saraogi, Devesh, et al.
Published: (2025)
Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows
by: Kumar, Shivani, et al.
Published: (2026)
by: Kumar, Shivani, et al.
Published: (2026)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
by: Peng, Xiangyu, et al.
Published: (2025)
by: Peng, Xiangyu, et al.
Published: (2025)
WorkTeam: Constructing Workflows from Natural Language with Multi-Agents
by: Liu, Hanchao, et al.
Published: (2025)
by: Liu, Hanchao, et al.
Published: (2025)
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI
by: Zhang, Jianguo, et al.
Published: (2023)
by: Zhang, Jianguo, et al.
Published: (2023)
Agent Workflow Memory
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Shared Imagination: LLMs Hallucinate Alike
by: Zhou, Yilun, et al.
Published: (2024)
by: Zhou, Yilun, et al.
Published: (2024)
Conversation AI Dialog for Medicare powered by Finetuning and Retrieval Augmented Generation
by: Agrawal, Atharva Mangeshkumar, et al.
Published: (2025)
by: Agrawal, Atharva Mangeshkumar, et al.
Published: (2025)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
by: Laban, Philippe, et al.
Published: (2023)
by: Laban, Philippe, et al.
Published: (2023)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
by: Laban, Philippe, et al.
Published: (2024)
by: Laban, Philippe, et al.
Published: (2024)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
by: Chen, Hardy, et al.
Published: (2026)
by: Chen, Hardy, et al.
Published: (2026)
Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution
by: Cifani, Susanna, et al.
Published: (2026)
by: Cifani, Susanna, et al.
Published: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
by: Kartik, NVJK, et al.
Published: (2025)
by: Kartik, NVJK, et al.
Published: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
by: Cao, Ruisheng, et al.
Published: (2024)
by: Cao, Ruisheng, et al.
Published: (2024)
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
by: Park, Chiwan, et al.
Published: (2025)
by: Park, Chiwan, et al.
Published: (2025)
Composable NLP Workflows for BERT-based Ranking and QA System
by: Kumar, Gaurav, et al.
Published: (2025)
by: Kumar, Gaurav, et al.
Published: (2025)
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
by: Lei, Fangyu, et al.
Published: (2024)
by: Lei, Fangyu, et al.
Published: (2024)
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
by: Grolleau, François, et al.
Published: (2025)
by: Grolleau, François, et al.
Published: (2025)
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
by: Chen, Haolin, et al.
Published: (2026)
by: Chen, Haolin, et al.
Published: (2026)
Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows
by: Sebe, Clémence, et al.
Published: (2024)
by: Sebe, Clémence, et al.
Published: (2024)
CreAgentive: An Agent Workflow Driven Multi-Category Creative Generation Engine
by: Cheng, Yuyang, et al.
Published: (2025)
by: Cheng, Yuyang, et al.
Published: (2025)
Similar Items
-
Benchmarking Deep Search over Heterogeneous Enterprise Data
by: Choubey, Prafulla Kumar, et al.
Published: (2025) -
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024) -
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024) -
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026) -
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)