FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Yitao, Hu, Tiansheng, Zhao, Yilun, Cohan, Arman, Zhao, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025)
by: Hu, Tiansheng, et al.
Published: (2025)
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
by: Zhao, Yilun, et al.
Published: (2024)
by: Zhao, Yilun, et al.
Published: (2024)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
by: Ding, Hang, et al.
Published: (2025)
by: Ding, Hang, et al.
Published: (2025)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
by: Hu, Yunhai, et al.
Published: (2025)
by: Hu, Yunhai, et al.
Published: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
FinTextQA: A Dataset for Long-form Financial Question Answering
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025)
by: Xu, Zhijian, et al.
Published: (2025)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
by: Shen, Yifei, et al.
Published: (2026)
by: Shen, Yifei, et al.
Published: (2026)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
LFQA-E: Carefully Benchmarking Long-form QA Evaluation
by: Fan, Yuchen, et al.
Published: (2024)
by: Fan, Yuchen, et al.
Published: (2024)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
by: Zhang, Siyue, et al.
Published: (2025)
by: Zhang, Siyue, et al.
Published: (2025)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
by: Wang, Chengye, et al.
Published: (2025)
by: Wang, Chengye, et al.
Published: (2025)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
by: Phanse, Rohan, et al.
Published: (2025)
by: Phanse, Rohan, et al.
Published: (2025)
Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
What Factors Affect LLMs and RLLMs in Financial Question Answering?
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Understanding Retrieval Augmentation for Long-Form Question Answering
by: Chen, Hung-Ting, et al.
Published: (2023)
by: Chen, Hung-Ting, et al.
Published: (2023)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
by: Li, Chuhan, et al.
Published: (2024)
by: Li, Chuhan, et al.
Published: (2024)
Improving Attributed Long-form Question Answering with Intent Awareness
by: Zhao, Xinran, et al.
Published: (2026)
by: Zhao, Xinran, et al.
Published: (2026)
Z1: Efficient Test-time Scaling with Code
by: Yu, Zhaojian, et al.
Published: (2025)
by: Yu, Zhaojian, et al.
Published: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
by: Deng, Chunyuan, et al.
Published: (2023)
by: Deng, Chunyuan, et al.
Published: (2023)
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
Atomic Consistency Preference Optimization for Long-Form Question Answering
by: Chen, Jingfeng, et al.
Published: (2025)
by: Chen, Jingfeng, et al.
Published: (2025)
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
by: Tang, Xiangru, et al.
Published: (2023)
by: Tang, Xiangru, et al.
Published: (2023)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge
by: He, Jie, et al.
Published: (2024)
by: He, Jie, et al.
Published: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
SciMDR: Advancing Scientific Multimodal Document Reasoning
by: Chen, Ziyu, et al.
Published: (2026)
by: Chen, Ziyu, et al.
Published: (2026)
Multi-Document Financial Question Answering using LLMs
by: Shah, Shalin, et al.
Published: (2024)
by: Shah, Shalin, et al.
Published: (2024)
Similar Items
-
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025) -
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
by: Zhao, Yilun, et al.
Published: (2024) -
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026) -
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
by: Ding, Hang, et al.
Published: (2025) -
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
by: Hu, Yunhai, et al.
Published: (2025)