LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
Fuente:
arXiv
Saved in:
| Main Authors: | Bishop, Jennifer A, Xie, Qianqian, Ananiadou, Sophia |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
FETILDA: An Effective Framework For Fin-tuned Embeddings For Long Financial Text Documents
by: Xia, Bolun "Namir", et al.
Published: (2022)
by: Xia, Bolun "Namir", et al.
Published: (2022)
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024)
by: Assogba, Yannick, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation
by: Li, Changmao, et al.
Published: (2024)
by: Li, Changmao, et al.
Published: (2024)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
by: Nadali, Alireza, et al.
Published: (2026)
by: Nadali, Alireza, et al.
Published: (2026)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
CSTRL: Context-Driven Sequential Transfer Learning for Abstractive Radiology Report Summarization
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2025)
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2025)
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
by: Günther, Michael, et al.
Published: (2023)
by: Günther, Michael, et al.
Published: (2023)
Observations on Building RAG Systems for Technical Documents
by: Soman, Sumit, et al.
Published: (2024)
by: Soman, Sumit, et al.
Published: (2024)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
by: Iqbal, Hasan, et al.
Published: (2024)
by: Iqbal, Hasan, et al.
Published: (2024)
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
by: Gao, Heyang, et al.
Published: (2025)
by: Gao, Heyang, et al.
Published: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
by: González-Pizarro, Felipe, et al.
Published: (2024)
by: González-Pizarro, Felipe, et al.
Published: (2024)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
by: Cha, Jungyoub, et al.
Published: (2025)
by: Cha, Jungyoub, et al.
Published: (2025)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
by: Dhole, Kaustubh
Published: (2025)
by: Dhole, Kaustubh
Published: (2025)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
by: Gilhuly, Colleen, et al.
Published: (2025)
by: Gilhuly, Colleen, et al.
Published: (2025)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
by: Zhou, Xiaoling, et al.
Published: (2024)
by: Zhou, Xiaoling, et al.
Published: (2024)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
Simple and Effective Baselines for Code Summarisation Evaluation
by: Robinson, Jade, et al.
Published: (2025)
by: Robinson, Jade, et al.
Published: (2025)
On Preserving the Knowledge of Long Clinical Texts
by: Hasan, Mohammad Junayed, et al.
Published: (2023)
by: Hasan, Mohammad Junayed, et al.
Published: (2023)
Dealing with Annotator Disagreement in Hate Speech Classification
by: Dehghan, Somaiyeh, et al.
Published: (2025)
by: Dehghan, Somaiyeh, et al.
Published: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
by: Hong, Chunsan, et al.
Published: (2025)
by: Hong, Chunsan, et al.
Published: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
Action-Item-Driven Summarization of Long Meeting Transcripts
by: Golia, Logan, et al.
Published: (2023)
by: Golia, Logan, et al.
Published: (2023)
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
by: Bandarkar, Lucas, et al.
Published: (2023)
by: Bandarkar, Lucas, et al.
Published: (2023)
DynaSemble: Dynamic Ensembling of Textual and Structure-Based Models for Knowledge Graph Completion
by: Nandi, Ananjan, et al.
Published: (2023)
by: Nandi, Ananjan, et al.
Published: (2023)
DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
by: Simon, Sona Elza, et al.
Published: (2025)
by: Simon, Sona Elza, et al.
Published: (2025)
Knowledge Graph Embeddings: A Comprehensive Survey on Capturing Relation Properties
by: Niu, Guanglin
Published: (2024)
by: Niu, Guanglin
Published: (2024)
Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realistic Coaching Agent Interactions
by: Yun, Taedong, et al.
Published: (2025)
by: Yun, Taedong, et al.
Published: (2025)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)
by: Oliveira, Rafael C. T.
Published: (2026)
Flash Multi-Head Feed-Forward Network
by: Zhang, Minshen, et al.
Published: (2025)
by: Zhang, Minshen, et al.
Published: (2025)
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
by: Sirbu, Iustin, et al.
Published: (2025)
by: Sirbu, Iustin, et al.
Published: (2025)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
by: Gao, Yutong, et al.
Published: (2026)
by: Gao, Yutong, et al.
Published: (2026)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
by: Yin, Yuwei, et al.
Published: (2026)
by: Yin, Yuwei, et al.
Published: (2026)
Similar Items
-
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025) -
FETILDA: An Effective Framework For Fin-tuned Embeddings For Long Financial Text Documents
by: Xia, Bolun "Namir", et al.
Published: (2022) -
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023) -
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation
by: Li, Changmao, et al.
Published: (2024)