Gespeichert in:
| Hauptverfasser: | Zheng, Danna, Lapata, Mirella, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.15792 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
von: Li, Miao, et al.
Veröffentlicht: (2026)
von: Li, Miao, et al.
Veröffentlicht: (2026)
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
von: Papoudakis, Argyrios, et al.
Veröffentlicht: (2026)
von: Papoudakis, Argyrios, et al.
Veröffentlicht: (2026)
BookWorm: A Dataset for Character Description and Analysis
von: Papoudakis, Argyrios, et al.
Veröffentlicht: (2024)
von: Papoudakis, Argyrios, et al.
Veröffentlicht: (2024)
Explanatory Summarization with Discourse-Driven Planning
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
von: Li, Junliang, et al.
Veröffentlicht: (2025)
von: Li, Junliang, et al.
Veröffentlicht: (2025)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2023)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2023)
Reasoning about Intent for Ambiguous Requests
von: Saparina, Irina, et al.
Veröffentlicht: (2025)
von: Saparina, Irina, et al.
Veröffentlicht: (2025)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
von: Adhikari, Ashutosh, et al.
Veröffentlicht: (2025)
von: Adhikari, Ashutosh, et al.
Veröffentlicht: (2025)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
von: Saparina, Irina, et al.
Veröffentlicht: (2025)
von: Saparina, Irina, et al.
Veröffentlicht: (2025)
K*-Means: A Parameter-free Clustering Algorithm
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
How Does Response Length Affect Long-Form Factuality
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Generating Visual Stories with Grounded and Coreferent Characters
von: Liu, Danyang, et al.
Veröffentlicht: (2024)
von: Liu, Danyang, et al.
Veröffentlicht: (2024)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2026)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2026)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
von: Yu, Simon, et al.
Veröffentlicht: (2023)
von: Yu, Simon, et al.
Veröffentlicht: (2023)
Learning to Reason for Long-Form Story Generation
von: Gurung, Alexander, et al.
Veröffentlicht: (2025)
von: Gurung, Alexander, et al.
Veröffentlicht: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
CHIRON: Rich Character Representations in Long-Form Narratives
von: Gurung, Alexander, et al.
Veröffentlicht: (2024)
von: Gurung, Alexander, et al.
Veröffentlicht: (2024)
Linguistic Calibration of Long-Form Generations
von: Band, Neil, et al.
Veröffentlicht: (2024)
von: Band, Neil, et al.
Veröffentlicht: (2024)
FactLens: Benchmarking Fine-Grained Fact Verification
von: Mitra, Kushan, et al.
Veröffentlicht: (2024)
von: Mitra, Kushan, et al.
Veröffentlicht: (2024)
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hu, Zhiwei, et al.
Veröffentlicht: (2024)
Are Large Language Models Table-based Fact-Checkers?
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
LongForm: Effective Instruction Tuning with Reverse Instructions
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
von: Opsahl, Tobias A.
Veröffentlicht: (2024)
von: Opsahl, Tobias A.
Veröffentlicht: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Fact-Consistency Evaluation of Text-to-SQL Generation for Business Intelligence Using Exaone 3.5
von: Choi, Jeho
Veröffentlicht: (2025)
von: Choi, Jeho
Veröffentlicht: (2025)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
von: Sun, Zhiqing, et al.
Veröffentlicht: (2024)
von: Sun, Zhiqing, et al.
Veröffentlicht: (2024)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
von: Fan, Haozhi, et al.
Veröffentlicht: (2026)
von: Fan, Haozhi, et al.
Veröffentlicht: (2026)
SimLM: Can Language Models Infer Parameters of Physical Systems?
von: Memery, Sean, et al.
Veröffentlicht: (2023)
von: Memery, Sean, et al.
Veröffentlicht: (2023)
InstructIE: A Bilingual Instruction-based Information Extraction Dataset
von: Gui, Honghao, et al.
Veröffentlicht: (2023)
von: Gui, Honghao, et al.
Veröffentlicht: (2023)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
von: Zheng, Danna, et al.
Veröffentlicht: (2024) -
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
von: Li, Miao, et al.
Veröffentlicht: (2026) -
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
von: Zheng, Danna, et al.
Veröffentlicht: (2024) -
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
von: Zheng, Danna, et al.
Veröffentlicht: (2024) -
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
von: Papoudakis, Argyrios, et al.
Veröffentlicht: (2026)