Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Bo, Yuan, Han, Pandelea, Vlad, Luo, Wuqiong, Zhao, Yingzhu, Ma, Zheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
par: Yuan, Han, et autres
Publié: (2025)
par: Yuan, Han, et autres
Publié: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
par: Long, Yitao, et autres
Publié: (2025)
par: Long, Yitao, et autres
Publié: (2025)
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering
par: Zhao, Qingfei, et autres
Publié: (2024)
par: Zhao, Qingfei, et autres
Publié: (2024)
FusionMind -- Improving question and answering with external context fusion
par: Verma, Shreyas, et autres
Publié: (2023)
par: Verma, Shreyas, et autres
Publié: (2023)
Long-context Non-factoid Question Answering in Indic Languages
par: Mishra, Ritwik, et autres
Publié: (2025)
par: Mishra, Ritwik, et autres
Publié: (2025)
LLM-based Triplet Extraction from Financial Reports
par: Wesslund, Dante, et autres
Publié: (2026)
par: Wesslund, Dante, et autres
Publié: (2026)
Which questions should I answer? Salience Prediction of Inquisitive Questions
par: Wu, Yating, et autres
Publié: (2024)
par: Wu, Yating, et autres
Publié: (2024)
Training data generation for context-dependent rubric-based short answer grading
par: Šindelář, Pavel, et autres
Publié: (2026)
par: Šindelář, Pavel, et autres
Publié: (2026)
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
par: Wu, Junjie, et autres
Publié: (2025)
par: Wu, Junjie, et autres
Publié: (2025)
ConSens: Assessing context grounding in open-book question answering
par: Vankov, Ivan, et autres
Publié: (2025)
par: Vankov, Ivan, et autres
Publié: (2025)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
par: Gupta, Lavanya, et autres
Publié: (2024)
par: Gupta, Lavanya, et autres
Publié: (2024)
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
par: Du, Haowei, et autres
Publié: (2024)
par: Du, Haowei, et autres
Publié: (2024)
ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering
par: Li, Huayang, et autres
Publié: (2024)
par: Li, Huayang, et autres
Publié: (2024)
Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering
par: Srivastava, Pragya, et autres
Publié: (2024)
par: Srivastava, Pragya, et autres
Publié: (2024)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
par: Kobeissi, Amine, et autres
Publié: (2026)
par: Kobeissi, Amine, et autres
Publié: (2026)
Question answering system of bridge design specification based on large language model
par: Zhang, Leye, et autres
Publié: (2024)
par: Zhang, Leye, et autres
Publié: (2024)
Question answering systems for health professionals at the point of care -- a systematic review
par: Kell, Gregory, et autres
Publié: (2024)
par: Kell, Gregory, et autres
Publié: (2024)
Evaluating Long-Term Memory for Long-Context Question Answering
par: Terranova, Alessandra, et autres
Publié: (2025)
par: Terranova, Alessandra, et autres
Publié: (2025)
FinTextQA: A Dataset for Long-form Financial Question Answering
par: Chen, Jian, et autres
Publié: (2024)
par: Chen, Jian, et autres
Publié: (2024)
What Factors Affect LLMs and RLLMs in Financial Question Answering?
par: Wang, Peng, et autres
Publié: (2025)
par: Wang, Peng, et autres
Publié: (2025)
LongGenBench: Long-context Generation Benchmark
par: Liu, Xiang, et autres
Publié: (2024)
par: Liu, Xiang, et autres
Publié: (2024)
Enhancing Financial Question Answering with a Multi-Agent Reflection Framework
par: Fatemi, Sorouralsadat, et autres
Publié: (2024)
par: Fatemi, Sorouralsadat, et autres
Publié: (2024)
New Evaluation Paradigm for Lexical Simplification
par: Qiang, Jipeng, et autres
Publié: (2025)
par: Qiang, Jipeng, et autres
Publié: (2025)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
par: Bachyr, Omar El, et autres
Publié: (2026)
par: Bachyr, Omar El, et autres
Publié: (2026)
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
par: Liu, Xinyu, et autres
Publié: (2024)
par: Liu, Xinyu, et autres
Publié: (2024)
Triplet-Block Diffusion RWKV
par: Lin, Ke, et autres
Publié: (2026)
par: Lin, Ke, et autres
Publié: (2026)
Long-context LLMs Struggle with Long In-context Learning
par: Li, Tianle, et autres
Publié: (2024)
par: Li, Tianle, et autres
Publié: (2024)
On Mechanistic Circuits for Extractive Question-Answering
par: Basu, Samyadeep, et autres
Publié: (2025)
par: Basu, Samyadeep, et autres
Publié: (2025)
DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature
par: Li, Dawei, et autres
Publié: (2024)
par: Li, Dawei, et autres
Publié: (2024)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
par: Jiang, Ziyan, et autres
Publié: (2024)
par: Jiang, Ziyan, et autres
Publié: (2024)
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
par: Wang, Renxi, et autres
Publié: (2025)
par: Wang, Renxi, et autres
Publié: (2025)
Are Large Language Models Good In-context Learners for Financial Sentiment Analysis?
par: Wei, Xinyu, et autres
Publié: (2025)
par: Wei, Xinyu, et autres
Publié: (2025)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
par: Wu, Siwei, et autres
Publié: (2025)
par: Wu, Siwei, et autres
Publié: (2025)
Asking and Answering Questions to Extract Event-Argument Structures
par: Uddin, Md Nayem, et autres
Publié: (2024)
par: Uddin, Md Nayem, et autres
Publié: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
par: Molfese, Francesco Maria, et autres
Publié: (2025)
par: Molfese, Francesco Maria, et autres
Publié: (2025)
Inter-Passage Verification for Multi-evidence Multi-answer QA
par: Chen, Bingsen, et autres
Publié: (2025)
par: Chen, Bingsen, et autres
Publié: (2025)
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
par: Gavin, Shawn, et autres
Publié: (2024)
par: Gavin, Shawn, et autres
Publié: (2024)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
par: Hu, Mengkang, et autres
Publié: (2024)
par: Hu, Mengkang, et autres
Publié: (2024)
GraphCheck: Breaking Long-Term Text Barriers with Extracted Knowledge Graph-Powered Fact-Checking
par: Chen, Yingjian, et autres
Publié: (2025)
par: Chen, Yingjian, et autres
Publié: (2025)
A New Sentence Extraction Strategy for Unsupervised Extractive Summarization Methods
par: Tao, Dehao, et autres
Publié: (2021)
par: Tao, Dehao, et autres
Publié: (2021)
Documents similaires
-
Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
par: Yuan, Han, et autres
Publié: (2025) -
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
par: Long, Yitao, et autres
Publié: (2025) -
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering
par: Zhao, Qingfei, et autres
Publié: (2024) -
FusionMind -- Improving question and answering with external context fusion
par: Verma, Shreyas, et autres
Publié: (2023) -
Long-context Non-factoid Question Answering in Indic Languages
par: Mishra, Ritwik, et autres
Publié: (2025)