Long-Form Information Alignment Evaluation Beyond Atomic Facts
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Danna, Lapata, Mirella, Pan, Jeff Z. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
di: Zheng, Danna, et al.
Pubblicazione: (2024)
di: Zheng, Danna, et al.
Pubblicazione: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
di: Li, Miao, et al.
Pubblicazione: (2026)
di: Li, Miao, et al.
Pubblicazione: (2026)
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
di: Zheng, Danna, et al.
Pubblicazione: (2024)
di: Zheng, Danna, et al.
Pubblicazione: (2024)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
di: Li, Junliang, et al.
Pubblicazione: (2025)
di: Li, Junliang, et al.
Pubblicazione: (2025)
Explanatory Summarization with Discourse-Driven Planning
di: Liu, Dongqi, et al.
Pubblicazione: (2025)
di: Liu, Dongqi, et al.
Pubblicazione: (2025)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
di: Papoudakis, Argyrios, et al.
Pubblicazione: (2026)
di: Papoudakis, Argyrios, et al.
Pubblicazione: (2026)
BookWorm: A Dataset for Character Description and Analysis
di: Papoudakis, Argyrios, et al.
Pubblicazione: (2024)
di: Papoudakis, Argyrios, et al.
Pubblicazione: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
di: Zheng, Danna, et al.
Pubblicazione: (2024)
di: Zheng, Danna, et al.
Pubblicazione: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
di: Gupta, Akash, et al.
Pubblicazione: (2025)
di: Gupta, Akash, et al.
Pubblicazione: (2025)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2023)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2023)
Reasoning about Intent for Ambiguous Requests
di: Saparina, Irina, et al.
Pubblicazione: (2025)
di: Saparina, Irina, et al.
Pubblicazione: (2025)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
di: Adhikari, Ashutosh, et al.
Pubblicazione: (2025)
di: Adhikari, Ashutosh, et al.
Pubblicazione: (2025)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
di: Saparina, Irina, et al.
Pubblicazione: (2025)
di: Saparina, Irina, et al.
Pubblicazione: (2025)
How Does Response Length Affect Long-Form Factuality
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
K*-Means: A Parameter-free Clustering Algorithm
di: Mahon, Louis, et al.
Pubblicazione: (2025)
di: Mahon, Louis, et al.
Pubblicazione: (2025)
Linguistic Calibration of Long-Form Generations
di: Band, Neil, et al.
Pubblicazione: (2024)
di: Band, Neil, et al.
Pubblicazione: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
di: Yu, Simon, et al.
Pubblicazione: (2023)
di: Yu, Simon, et al.
Pubblicazione: (2023)
Generating Visual Stories with Grounded and Coreferent Characters
di: Liu, Danyang, et al.
Pubblicazione: (2024)
di: Liu, Danyang, et al.
Pubblicazione: (2024)
FactLens: Benchmarking Fine-Grained Fact Verification
di: Mitra, Kushan, et al.
Pubblicazione: (2024)
di: Mitra, Kushan, et al.
Pubblicazione: (2024)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
di: Eisenstein, Jacob, et al.
Pubblicazione: (2026)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2026)
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
di: Burger, Christopher, et al.
Pubblicazione: (2024)
di: Burger, Christopher, et al.
Pubblicazione: (2024)
Are Large Language Models Table-based Fact-Checkers?
di: Zhang, Hanwen, et al.
Pubblicazione: (2024)
di: Zhang, Hanwen, et al.
Pubblicazione: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
di: Köksal, Abdullatif, et al.
Pubblicazione: (2023)
di: Köksal, Abdullatif, et al.
Pubblicazione: (2023)
Learning to Reason for Long-Form Story Generation
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
di: Hu, Zhiwei, et al.
Pubblicazione: (2024)
di: Hu, Zhiwei, et al.
Pubblicazione: (2024)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
di: Ardestani, MohamamdJavad, et al.
Pubblicazione: (2025)
di: Ardestani, MohamamdJavad, et al.
Pubblicazione: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
di: Obeso, Oscar, et al.
Pubblicazione: (2025)
di: Obeso, Oscar, et al.
Pubblicazione: (2025)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
di: Opsahl, Tobias A.
Pubblicazione: (2024)
di: Opsahl, Tobias A.
Pubblicazione: (2024)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
di: Huang, Wenyu, et al.
Pubblicazione: (2024)
di: Huang, Wenyu, et al.
Pubblicazione: (2024)
CHIRON: Rich Character Representations in Long-Form Narratives
di: Gurung, Alexander, et al.
Pubblicazione: (2024)
di: Gurung, Alexander, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
Fact-Consistency Evaluation of Text-to-SQL Generation for Business Intelligence Using Exaone 3.5
di: Choi, Jeho
Pubblicazione: (2025)
di: Choi, Jeho
Pubblicazione: (2025)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
di: Fan, Haozhi, et al.
Pubblicazione: (2026)
di: Fan, Haozhi, et al.
Pubblicazione: (2026)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
di: Zhang, Yifan, et al.
Pubblicazione: (2024)
di: Zhang, Yifan, et al.
Pubblicazione: (2024)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
di: Zheng, Danna, et al.
Pubblicazione: (2024) -
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
di: Li, Miao, et al.
Pubblicazione: (2026) -
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
di: Zheng, Danna, et al.
Pubblicazione: (2024) -
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
di: Li, Junliang, et al.
Pubblicazione: (2025) -
Explanatory Summarization with Discourse-Driven Planning
di: Liu, Dongqi, et al.
Pubblicazione: (2025)