All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kulkarni, Shubham, Lyzhov, Alexander, Chaitanya, Shiva, Joshi, Preetam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection
von: Kulkarni, Shubham, et al.
Veröffentlicht: (2026)
von: Kulkarni, Shubham, et al.
Veröffentlicht: (2026)
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
von: Paudel, Bibek, et al.
Veröffentlicht: (2025)
von: Paudel, Bibek, et al.
Veröffentlicht: (2025)
Position: Human-Centric AI Requires a Minimum Viable Level of Human Understanding
von: Lin, Fangzhou, et al.
Veröffentlicht: (2026)
von: Lin, Fangzhou, et al.
Veröffentlicht: (2026)
Transformers are Graph Neural Networks
von: Joshi, Chaitanya K.
Veröffentlicht: (2025)
von: Joshi, Chaitanya K.
Veröffentlicht: (2025)
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
von: Vatsal, Shubham, et al.
Veröffentlicht: (2026)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2026)
MedAidDialog: A Multilingual Multi-Turn Medical Dialogue Dataset for Accessible Healthcare
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
AI Evaluation Should Require Standardized Item-Level Data Releases
von: Jiang, Han, et al.
Veröffentlicht: (2026)
von: Jiang, Han, et al.
Veröffentlicht: (2026)
Explainable Human-AI Interaction: A Planning Perspective
von: Sreedharan, Sarath, et al.
Veröffentlicht: (2024)
von: Sreedharan, Sarath, et al.
Veröffentlicht: (2024)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
Hybrid Neuro-Symbolic Models for Ethical AI in Risk-Sensitive Domains
von: Kolli, Chaitanya Kumar
Veröffentlicht: (2025)
von: Kolli, Chaitanya Kumar
Veröffentlicht: (2025)
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
Towards Dialogues for Joint Human-AI Reasoning and Value Alignment
von: Bezou-Vrakatseli, Elfia, et al.
Veröffentlicht: (2024)
von: Bezou-Vrakatseli, Elfia, et al.
Veröffentlicht: (2024)
Decision-Oriented Dialogue for Human-AI Collaboration
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
ChannelFlow-Tools: A Standardized Dataset Creation Pipeline for 3D Obstructed Channel Flows
von: Kavane, Shubham, et al.
Veröffentlicht: (2025)
von: Kavane, Shubham, et al.
Veröffentlicht: (2025)
Ethical Framework for Harnessing the Power of AI in Healthcare and Beyond
von: Nasir, Sidra, et al.
Veröffentlicht: (2023)
von: Nasir, Sidra, et al.
Veröffentlicht: (2023)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
von: Arias-Duart, Anna, et al.
Veröffentlicht: (2025)
von: Arias-Duart, Anna, et al.
Veröffentlicht: (2025)
Beyond Accuracy: A Decision-Theoretic Framework for Allocation-Aware Healthcare AI
von: Ferzana, Rifa
Veröffentlicht: (2026)
von: Ferzana, Rifa
Veröffentlicht: (2026)
Dialogue with the Machine and Dialogue with the Art World: Evaluating Generative AI for Culturally-Situated Creativity
von: Qadri, Rida, et al.
Veröffentlicht: (2024)
von: Qadri, Rida, et al.
Veröffentlicht: (2024)
Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI
von: Perera, Dilruk, et al.
Veröffentlicht: (2025)
von: Perera, Dilruk, et al.
Veröffentlicht: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
von: Shastry, KN Ajay, et al.
Veröffentlicht: (2026)
von: Shastry, KN Ajay, et al.
Veröffentlicht: (2026)
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials
von: Joshi, Chaitanya K., et al.
Veröffentlicht: (2025)
von: Joshi, Chaitanya K., et al.
Veröffentlicht: (2025)
Dynamic Evaluation Framework for Personalized and Trustworthy Agents: A Multi-Session Approach to Preference Adaptability
von: Shah, Chirag, et al.
Veröffentlicht: (2025)
von: Shah, Chirag, et al.
Veröffentlicht: (2025)
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations
von: Ebubechukwu, Ike, et al.
Veröffentlicht: (2024)
von: Ebubechukwu, Ike, et al.
Veröffentlicht: (2024)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
von: Zhu, Ruizhe, et al.
Veröffentlicht: (2025)
von: Zhu, Ruizhe, et al.
Veröffentlicht: (2025)
Empathy by Design: Aligning Large Language Models for Healthcare Dialogue
von: Umucu, Emre, et al.
Veröffentlicht: (2025)
von: Umucu, Emre, et al.
Veröffentlicht: (2025)
Beyond Levels of Driving Automation: A Triadic Framework of Human-AI Collaboration in On-Road Mobility
von: Huang, Gaojian, et al.
Veröffentlicht: (2025)
von: Huang, Gaojian, et al.
Veröffentlicht: (2025)
Defining Explainable AI for Requirements Analysis
von: Sheh, Raymond, et al.
Veröffentlicht: (2026)
von: Sheh, Raymond, et al.
Veröffentlicht: (2026)
JEDA: Query-Free Clinical Order Search from Ambient Dialogues
von: Singh, Praphul, et al.
Veröffentlicht: (2025)
von: Singh, Praphul, et al.
Veröffentlicht: (2025)
We Are All Creators: Generative AI, Collective Knowledge, and the Path Towards Human-AI Synergy
von: Linares-Pellicer, Jordi, et al.
Veröffentlicht: (2025)
von: Linares-Pellicer, Jordi, et al.
Veröffentlicht: (2025)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment
von: Levy, Oz, et al.
Veröffentlicht: (2026)
von: Levy, Oz, et al.
Veröffentlicht: (2026)
Towards Negotiative Dialogue for the Talkamatic Dialogue Manager
von: Larsson, Staffan, et al.
Veröffentlicht: (2024)
von: Larsson, Staffan, et al.
Veröffentlicht: (2024)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
von: Lu, Dongxu, et al.
Veröffentlicht: (2025)
von: Lu, Dongxu, et al.
Veröffentlicht: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
von: El, Batu, et al.
Veröffentlicht: (2025)
von: El, Batu, et al.
Veröffentlicht: (2025)
The Explanation Necessity for Healthcare AI
von: Mamalakis, Michail, et al.
Veröffentlicht: (2024)
von: Mamalakis, Michail, et al.
Veröffentlicht: (2024)
Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability
von: Zhu, Judy, et al.
Veröffentlicht: (2026)
von: Zhu, Judy, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
von: Su, Miao, et al.
Veröffentlicht: (2026)
von: Su, Miao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection
von: Kulkarni, Shubham, et al.
Veröffentlicht: (2026) -
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
von: Paudel, Bibek, et al.
Veröffentlicht: (2025) -
Position: Human-Centric AI Requires a Minimum Viable Level of Human Understanding
von: Lin, Fangzhou, et al.
Veröffentlicht: (2026) -
Transformers are Graph Neural Networks
von: Joshi, Chaitanya K.
Veröffentlicht: (2025) -
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
von: Vatsal, Shubham, et al.
Veröffentlicht: (2026)