On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
Fuente:
arXiv
Saved in:
| Main Authors: | Upadhyay, Shivani, Ataey, Messiah, Murtaza, Syed Shariyar, Nie, Yifan, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
Multi-Document Financial Question Answering using LLMs
by: Shah, Shalin, et al.
Published: (2024)
by: Shah, Shalin, et al.
Published: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025)
by: Pradeep, Ronak, et al.
Published: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Tracing Content Requirements in Financial Documents using Multi-granularity Text Analysis
by: Li, Xiaochen, et al.
Published: (2021)
by: Li, Xiaochen, et al.
Published: (2021)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
by: Jedidi, Nour, et al.
Published: (2026)
by: Jedidi, Nour, et al.
Published: (2026)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
by: Lin, Jimmy
Published: (2024)
by: Lin, Jimmy
Published: (2024)
Unifying Multimodal Retrieval via Document Screenshot Embedding
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
Leveraging LLMs to Evaluate Usefulness of Document
by: Wang, Xingzhu, et al.
Published: (2025)
by: Wang, Xingzhu, et al.
Published: (2025)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
A Multi-Task Embedder For Retrieval Augmented LLMs
by: Zhang, Peitian, et al.
Published: (2023)
by: Zhang, Peitian, et al.
Published: (2023)
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
by: Jin, Yichao, et al.
Published: (2025)
by: Jin, Yichao, et al.
Published: (2025)
Revisiting Feedback Models for HyDE
by: Jedidi, Nour, et al.
Published: (2025)
by: Jedidi, Nour, et al.
Published: (2025)
Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents
by: Sharifymoghaddam, Sahel, et al.
Published: (2026)
by: Sharifymoghaddam, Sahel, et al.
Published: (2026)
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
by: Zhuang, Shengyao, et al.
Published: (2024)
by: Zhuang, Shengyao, et al.
Published: (2024)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Tools are under-documented: Simple Document Expansion Boosts Tool Retrieval
by: Lu, Xuan, et al.
Published: (2025)
by: Lu, Xuan, et al.
Published: (2025)
Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
RankLLM: A Python Package for Reranking with LLMs
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
by: Tamber, Manveer Singh, et al.
Published: (2024)
by: Tamber, Manveer Singh, et al.
Published: (2024)
Musings About the Future of Search: A Return to the Past?
by: Lin, Jimmy, et al.
Published: (2024)
by: Lin, Jimmy, et al.
Published: (2024)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset
by: Rasool, Zafaryab, et al.
Published: (2023)
by: Rasool, Zafaryab, et al.
Published: (2023)
Optimizing Retrieval Strategies for Financial Question Answering Documents in Retrieval-Augmented Generation Systems
by: Kim, Sejong, et al.
Published: (2025)
by: Kim, Sejong, et al.
Published: (2025)
FinEmbedDiff: A Cost-Effective Approach of Classifying Financial Documents with Vector Sampling using Multi-modal Embedding Models
by: Biswas, Anjanava, et al.
Published: (2024)
by: Biswas, Anjanava, et al.
Published: (2024)
Weighted KL-Divergence for Document Ranking Model Refinement
by: Yang, Yingrui, et al.
Published: (2024)
by: Yang, Yingrui, et al.
Published: (2024)
Read the Docs Before Rewriting: Equip Rewriter with Domain Knowledge via Continual Pre-training
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
MealRec$^+$: A Meal Recommendation Dataset with Meal-Course Affiliation for Personalization and Healthiness
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Similar Items
-
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024) -
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
by: Sharifymoghaddam, Sahel, et al.
Published: (2024) -
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
by: Sharifymoghaddam, Sahel, et al.
Published: (2025) -
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
by: Upadhyay, Shivani, et al.
Published: (2024) -
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)