FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Magomere, Jabez, Kochkina, Elena, Mensah, Samuel, Kaur, Simerjot, Smiley, Charese H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025)
by: Mensah, Samuel, et al.
Published: (2025)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024)
by: Singh, Kuldeep, et al.
Published: (2024)
Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
by: Aguda, Toyin, et al.
Published: (2024)
by: Aguda, Toyin, et al.
Published: (2024)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)
by: Haque, Mirazul, et al.
Published: (2026)
Grounding LLM Reasoning with Knowledge Graphs
by: Amayuelas, Alfonso, et al.
Published: (2025)
by: Amayuelas, Alfonso, et al.
Published: (2025)
Conservative Bias in Large Language Models: Measuring Relation Predictions
by: Aguda, Toyin, et al.
Published: (2025)
by: Aguda, Toyin, et al.
Published: (2025)
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
by: Fons, Elizabeth, et al.
Published: (2025)
by: Fons, Elizabeth, et al.
Published: (2025)
How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
by: Khanmohammadi, Reza, et al.
Published: (2026)
by: Khanmohammadi, Reza, et al.
Published: (2026)
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
by: Khanmohammadi, Reza, et al.
Published: (2026)
by: Khanmohammadi, Reza, et al.
Published: (2026)
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
Calibrating LLM Confidence by Probing Perturbed Representation Stability
by: Khanmohammadi, Reza, et al.
Published: (2025)
by: Khanmohammadi, Reza, et al.
Published: (2025)
ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
MorphNLI: A Stepwise Approach to Natural Language Inference Using Text Morphing
by: Negru, Vlad Andrei, et al.
Published: (2025)
by: Negru, Vlad Andrei, et al.
Published: (2025)
MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference
by: Sadat, Mobashir, et al.
Published: (2024)
by: Sadat, Mobashir, et al.
Published: (2024)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
by: Choi, Chanyeol, et al.
Published: (2025)
by: Choi, Chanyeol, et al.
Published: (2025)
FinS-Pilot: A Benchmark for Online Financial RAG System
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
FinBen: A Holistic Financial Benchmark for Large Language Models
by: Xie, Qianqian, et al.
Published: (2024)
by: Xie, Qianqian, et al.
Published: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
by: Luo, Junyu, et al.
Published: (2025)
by: Luo, Junyu, et al.
Published: (2025)
FinTextQA: A Dataset for Long-form Financial Question Answering
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
by: Shah, Arjun, et al.
Published: (2024)
by: Shah, Arjun, et al.
Published: (2024)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
by: Kim, Eric Y., et al.
Published: (2026)
by: Kim, Eric Y., et al.
Published: (2026)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025)
by: Xie, Zhuohan, et al.
Published: (2025)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
Identifying Multiple Personalities in Large Language Models with External Evaluation
by: Song, Xiaoyang, et al.
Published: (2024)
by: Song, Xiaoyang, et al.
Published: (2024)
Synthetic Lyrics Detection Across Languages and Genres
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
FinLLM-B: When Large Language Models Meet Financial Breakout Trading
by: Zhang, Kang, et al.
Published: (2024)
by: Zhang, Kang, et al.
Published: (2024)
SocialNLI: A Dialogue-Centric Social Inference Dataset
by: Deo, Akhil, et al.
Published: (2025)
by: Deo, Akhil, et al.
Published: (2025)
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
by: Khanmohammadi, Reza, et al.
Published: (2024)
by: Khanmohammadi, Reza, et al.
Published: (2024)
Reverse-engineering NLI: A study of the meta-inferential properties of Natural Language Inference
by: Blanck, Rasmus, et al.
Published: (2026)
by: Blanck, Rasmus, et al.
Published: (2026)
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Similar Items
-
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025) -
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026) -
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024) -
Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
by: Aguda, Toyin, et al.
Published: (2024) -
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)