A Comparative Analysis of Faithfulness Metrics and Humans in Citation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Weijia, Aliannejadi, Mohammad, Pei, Jiahuan, Yuan, Yifei, Huang, Jia-Hong, Kanoulas, Evangelos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
Beyond Relevant Documents: A Knowledge-Intensive Approach for Query-Focused Summarization using Large Language Models
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering
by: Sidiropoulos, Georgios, et al.
Published: (2024)
by: Sidiropoulos, Georgios, et al.
Published: (2024)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
Data Augmentation for Conversational AI
by: Soudani, Heydar, et al.
Published: (2023)
by: Soudani, Heydar, et al.
Published: (2023)
Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
PSCon: Product Search Through Conversations
by: Zou, Jie, et al.
Published: (2025)
by: Zou, Jie, et al.
Published: (2025)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Generative Retrieval with Few-shot Indexing
by: Askari, Arian, et al.
Published: (2024)
by: Askari, Arian, et al.
Published: (2024)
Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
by: Petcu, Roxana, et al.
Published: (2025)
by: Petcu, Roxana, et al.
Published: (2025)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
IRLab@iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
SubSearch: Intermediate Rewards for Unsupervised Guided Reasoning in Complex Retrieval
by: Petcu, Roxana, et al.
Published: (2026)
by: Petcu, Roxana, et al.
Published: (2026)
Learning to Ask: Conversational Product Search via Representation Learning
by: Zou, Jie, et al.
Published: (2024)
by: Zou, Jie, et al.
Published: (2024)
Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrieval
by: Eichholtz, Arne, et al.
Published: (2026)
by: Eichholtz, Arne, et al.
Published: (2026)
A Survey on Recent Advances in Conversational Data Generation
by: Soudani, Heydar, et al.
Published: (2024)
by: Soudani, Heydar, et al.
Published: (2024)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Can We Use Large Language Models to Fill Relevance Judgment Holes?
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
by: Mousavian, Maryam, et al.
Published: (2025)
by: Mousavian, Maryam, et al.
Published: (2025)
Conversational Search: From Fundamentals to Frontiers in the LLM Era
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
by: Ramezan, Kimia, et al.
Published: (2025)
by: Ramezan, Kimia, et al.
Published: (2025)
TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
Investigating LLM Variability in Personalized Conversational Information Retrieval
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
Faithful Temporal Question Answering over Heterogeneous Sources
by: Jia, Zhen, et al.
Published: (2024)
by: Jia, Zhen, et al.
Published: (2024)
Diagnosing and Repairing Citation Failures in Generative Engine Optimization
by: Tian, Zhihua, et al.
Published: (2026)
by: Tian, Zhihua, et al.
Published: (2026)
Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification
by: Yuan, Yifei, et al.
Published: (2025)
by: Yuan, Yifei, et al.
Published: (2025)
A Novel Evaluation Framework for Image2Text Generation
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
by: Ma, Yongqiang, et al.
Published: (2024)
by: Ma, Yongqiang, et al.
Published: (2024)
TREC iKAT 2023: The Interactive Knowledge Assistance Track Overview
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025)
by: Hu, Caiyu, et al.
Published: (2025)
Improving the Robustness of Dense Retrievers Against Typos via Multi-Positive Contrastive Learning
by: Sidiropoulos, Georgios, et al.
Published: (2024)
by: Sidiropoulos, Georgios, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis
by: Soudani, Heydar, et al.
Published: (2025)
by: Soudani, Heydar, et al.
Published: (2025)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
RAC: Retrieval-Augmented Clarification for Faithful Conversational Search
by: Kebir, Ahmed Rayane, et al.
Published: (2026)
by: Kebir, Ahmed Rayane, et al.
Published: (2026)
Similar Items
-
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
by: Zhang, Weijia, et al.
Published: (2024) -
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
by: Lupart, Simon, et al.
Published: (2025) -
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
by: Lupart, Simon, et al.
Published: (2024) -
Beyond Relevant Documents: A Knowledge-Intensive Approach for Query-Focused Summarization using Large Language Models
by: Zhang, Weijia, et al.
Published: (2024) -
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering
by: Sidiropoulos, Georgios, et al.
Published: (2024)