Climate Finance Bench
Fuente:
arXiv
Saved in:
| Main Authors: | Mankour, Rafik, Chafai, Yassine, Saleh, Hamada, Hassine, Ghassen Ben, Barreau, Thibaud, Tankov, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
by: Afzoon, Saleh, et al.
Published: (2024)
by: Afzoon, Saleh, et al.
Published: (2024)
AI for Climate Finance: Agentic Retrieval and Multi-Step Reasoning for Early Warning System Investments
by: Vaghefi, Saeid Ario, et al.
Published: (2025)
by: Vaghefi, Saeid Ario, et al.
Published: (2025)
Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models
by: Huang, Yin Jou, et al.
Published: (2025)
by: Huang, Yin Jou, et al.
Published: (2025)
How Personality Traits Influence Negotiation Outcomes? A Simulation based on Large Language Models
by: Huang, Yin Jou, et al.
Published: (2024)
by: Huang, Yin Jou, et al.
Published: (2024)
Algorithm for Automatic Legislative Text Consolidation
by: Etcheverry, Matias, et al.
Published: (2025)
by: Etcheverry, Matias, et al.
Published: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?
by: Harada, Yuto, et al.
Published: (2026)
by: Harada, Yuto, et al.
Published: (2026)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
Transformer-Based Model for Multilingual Hope Speech Detection
by: Ashraf, Nsrin, et al.
Published: (2026)
by: Ashraf, Nsrin, et al.
Published: (2026)
CoverBench: A Challenging Benchmark for Complex Claim Verification
by: Jacovi, Alon, et al.
Published: (2024)
by: Jacovi, Alon, et al.
Published: (2024)
Self-supervised Analogical Learning using Language Models
by: Zhou, Ben, et al.
Published: (2025)
by: Zhou, Ben, et al.
Published: (2025)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
by: Benhenda, Mostapha
Published: (2026)
by: Benhenda, Mostapha
Published: (2026)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
by: Ma, Congbo, et al.
Published: (2026)
by: Ma, Congbo, et al.
Published: (2026)
dzStance at StanceEval2024: Arabic Stance Detection based on Sentence Transformers
by: Lichouri, Mohamed, et al.
Published: (2024)
by: Lichouri, Mohamed, et al.
Published: (2024)
Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt
by: Anderson, Peter, et al.
Published: (2024)
by: Anderson, Peter, et al.
Published: (2024)
Role-Playing Evaluation for Large Language Models
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
BIG-Bench Extra Hard
by: Kazemi, Mehran, et al.
Published: (2025)
by: Kazemi, Mehran, et al.
Published: (2025)
Language-Aware Information Maximization for Transductive Few-Shot CLIP
by: Baklouti, Ghassen, et al.
Published: (2025)
by: Baklouti, Ghassen, et al.
Published: (2025)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
by: Gan, Ziliang, et al.
Published: (2024)
by: Gan, Ziliang, et al.
Published: (2024)
An investigation of structures responsible for gender bias in BERT and DistilBERT
by: Leteno, Thibaud, et al.
Published: (2024)
by: Leteno, Thibaud, et al.
Published: (2024)
AbsenceBench: Language Models Can't Tell What's Missing
by: Fu, Harvey Yiyun, et al.
Published: (2025)
by: Fu, Harvey Yiyun, et al.
Published: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
Revolutionizing Finance with LLMs: An Overview of Applications and Insights
by: Zhao, Huaqin, et al.
Published: (2024)
by: Zhao, Huaqin, et al.
Published: (2024)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon
by: Smaili, Kamel, et al.
Published: (2026)
by: Smaili, Kamel, et al.
Published: (2026)
Ebisu: Benchmarking Large Language Models in Japanese Finance
by: Peng, Xueqing, et al.
Published: (2026)
by: Peng, Xueqing, et al.
Published: (2026)
Text2Model: Generating dynamic chemical reactor models using large language models (LLMs)
by: Rupprecht, Sophia, et al.
Published: (2025)
by: Rupprecht, Sophia, et al.
Published: (2025)
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?
by: Chawla, Divij, et al.
Published: (2025)
by: Chawla, Divij, et al.
Published: (2025)
Expect the Unexpected: FailSafe Long Context QA for Finance
by: Kamble, Kiran, et al.
Published: (2025)
by: Kamble, Kiran, et al.
Published: (2025)
Fair Text Classification via Transferable Representations
by: Leteno, Thibaud, et al.
Published: (2025)
by: Leteno, Thibaud, et al.
Published: (2025)
Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance
by: Peng, Xueqing, et al.
Published: (2025)
by: Peng, Xueqing, et al.
Published: (2025)
Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance
by: Qian, Lingfei, et al.
Published: (2025)
by: Qian, Lingfei, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
by: Jia, Haoxuan, et al.
Published: (2026)
by: Jia, Haoxuan, et al.
Published: (2026)
Context-Masked Meta-Prompting for Privacy-Preserving LLM Adaptation in Finance
by: Hiraou, Sayash Raaj
Published: (2024)
by: Hiraou, Sayash Raaj
Published: (2024)
'Finance Wizard' at the FinLLM Challenge Task: Financial Text Summarization
by: Lee, Meisin, et al.
Published: (2024)
by: Lee, Meisin, et al.
Published: (2024)
Similar Items
-
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023) -
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
by: Afzoon, Saleh, et al.
Published: (2024) -
AI for Climate Finance: Agentic Retrieval and Multi-Step Reasoning for Early Warning System Investments
by: Vaghefi, Saeid Ario, et al.
Published: (2025) -
Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models
by: Huang, Yin Jou, et al.
Published: (2025) -
How Personality Traits Influence Negotiation Outcomes? A Simulation based on Large Language Models
by: Huang, Yin Jou, et al.
Published: (2024)