Salvato in:
| Autori principali: | Mayilvaghanan, Kawin, Gupta, Siddhant, Kumar, Ayush |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.14970 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries
di: Mayilvaghanan, Kawin, et al.
Pubblicazione: (2025)
di: Mayilvaghanan, Kawin, et al.
Pubblicazione: (2025)
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
di: Nathan, Varun, et al.
Pubblicazione: (2026)
di: Nathan, Varun, et al.
Pubblicazione: (2026)
Integration of LLM Quality Assurance into an NLG System
di: Chen, Ching-Yi, et al.
Pubblicazione: (2025)
di: Chen, Ching-Yi, et al.
Pubblicazione: (2025)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
Counterfactual Graph for Multi-Agent LLM Calibration
di: Huang, Jiatan, et al.
Pubblicazione: (2026)
di: Huang, Jiatan, et al.
Pubblicazione: (2026)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
di: Liu, Lijia, et al.
Pubblicazione: (2025)
di: Liu, Lijia, et al.
Pubblicazione: (2025)
LLM-Based Insight Extraction for Contact Center Analytics and Cost-Efficient Deployment
di: Embar, Varsha, et al.
Pubblicazione: (2025)
di: Embar, Varsha, et al.
Pubblicazione: (2025)
Multi-Facet Counterfactual Learning for Content Quality Evaluation
di: Zheng, Jiasheng, et al.
Pubblicazione: (2024)
di: Zheng, Jiasheng, et al.
Pubblicazione: (2024)
Equal Access, Unequal Interaction: A Counterfactual Audit of LLM Fairness
di: Amiri-Margavi, Alireza, et al.
Pubblicazione: (2026)
di: Amiri-Margavi, Alireza, et al.
Pubblicazione: (2026)
Uncertainty and Fairness Awareness in LLM-Based Recommendation Systems
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2026)
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2026)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
Aligning (Medical) LLMs for (Counterfactual) Fairness
di: Poulain, Raphael, et al.
Pubblicazione: (2024)
di: Poulain, Raphael, et al.
Pubblicazione: (2024)
Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG
di: Khadilkar, Harshad, et al.
Pubblicazione: (2025)
di: Khadilkar, Harshad, et al.
Pubblicazione: (2025)
Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies
di: Kulkarni, Siddhant, et al.
Pubblicazione: (2026)
di: Kulkarni, Siddhant, et al.
Pubblicazione: (2026)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
di: Wu, Wanxing, et al.
Pubblicazione: (2026)
di: Wu, Wanxing, et al.
Pubblicazione: (2026)
Agent-as-a-Graph: Knowledge Graph-Based Tool and Agent Retrieval for LLM Multi-Agent Systems
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
Question Answering on Patient Medical Records with Private Fine-Tuned LLMs
di: Kothari, Sara, et al.
Pubblicazione: (2025)
di: Kothari, Sara, et al.
Pubblicazione: (2025)
IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages
di: Gupta, Siddhant, et al.
Pubblicazione: (2024)
di: Gupta, Siddhant, et al.
Pubblicazione: (2024)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
Evaluating the Retrieval Component in LLM-Based Question Answering Systems
di: Alinejad, Ashkan, et al.
Pubblicazione: (2024)
di: Alinejad, Ashkan, et al.
Pubblicazione: (2024)
Anchor Points: Benchmarking Models with Much Fewer Examples
di: Vivek, Rajan, et al.
Pubblicazione: (2023)
di: Vivek, Rajan, et al.
Pubblicazione: (2023)
Stability Analysis of ChatGPT-based Sentiment Analysis in AI Quality Assurance
di: Ouyang, Tinghui, et al.
Pubblicazione: (2024)
di: Ouyang, Tinghui, et al.
Pubblicazione: (2024)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
di: Mo, Kaijie, et al.
Pubblicazione: (2026)
di: Mo, Kaijie, et al.
Pubblicazione: (2026)
Introducing Super RAGs in Mistral 8x7B-v1
di: Thakur, Ayush, et al.
Pubblicazione: (2024)
di: Thakur, Ayush, et al.
Pubblicazione: (2024)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
di: Baumgärtner, Tim, et al.
Pubblicazione: (2026)
di: Baumgärtner, Tim, et al.
Pubblicazione: (2026)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
di: Nawale, Janki Atul, et al.
Pubblicazione: (2025)
di: Nawale, Janki Atul, et al.
Pubblicazione: (2025)
Text-Based Detection of On-Hold Scripts in Contact Center Calls
di: Galimzianov, Dmitrii, et al.
Pubblicazione: (2024)
di: Galimzianov, Dmitrii, et al.
Pubblicazione: (2024)
LLM-Based Support for Diabetes Diagnosis: Opportunities, Scenarios, and Challenges with GPT-5
di: Gupta, Gaurav Kumar, et al.
Pubblicazione: (2025)
di: Gupta, Gaurav Kumar, et al.
Pubblicazione: (2025)
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
di: Khatchadourian, Raffi
Pubblicazione: (2026)
di: Khatchadourian, Raffi
Pubblicazione: (2026)
LLM-Based Section Identifiers Excel on Open Source but Stumble in Real World Applications
di: Krishnamoorthy, Saranya, et al.
Pubblicazione: (2024)
di: Krishnamoorthy, Saranya, et al.
Pubblicazione: (2024)
Tool-to-Agent Retrieval: Bridging Tools and Agents for Scalable LLM Multi-Agent Systems
di: Lumer, Elias, et al.
Pubblicazione: (2025)
di: Lumer, Elias, et al.
Pubblicazione: (2025)
Data Checklist: On Unit-Testing Datasets with Usable Information
di: Zhang, Heidi C., et al.
Pubblicazione: (2024)
di: Zhang, Heidi C., et al.
Pubblicazione: (2024)
Do Cognitively Interpretable Reasoning Traces Improve LLM Performance?
di: Bhambri, Siddhant, et al.
Pubblicazione: (2025)
di: Bhambri, Siddhant, et al.
Pubblicazione: (2025)
Substance over Style: Evaluating Proactive Conversational Coaching Agents
di: Srinivas, Vidya, et al.
Pubblicazione: (2025)
di: Srinivas, Vidya, et al.
Pubblicazione: (2025)
FairFlow: An Automated Approach to Model-based Counterfactual Data Augmentation For NLP
di: Tokpo, Ewoenam Kwaku, et al.
Pubblicazione: (2024)
di: Tokpo, Ewoenam Kwaku, et al.
Pubblicazione: (2024)
Are VLMs Really Blind
di: Singh, Ayush, et al.
Pubblicazione: (2024)
di: Singh, Ayush, et al.
Pubblicazione: (2024)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
di: Hillebrand, Lars, et al.
Pubblicazione: (2025)
di: Hillebrand, Lars, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries
di: Mayilvaghanan, Kawin, et al.
Pubblicazione: (2025) -
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
di: Nathan, Varun, et al.
Pubblicazione: (2026) -
Integration of LLM Quality Assurance into an NLG System
di: Chen, Ching-Yi, et al.
Pubblicazione: (2025) -
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025) -
Counterfactual Graph for Multi-Agent LLM Calibration
di: Huang, Jiatan, et al.
Pubblicazione: (2026)