ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | He, Chaoyue, Zhou, Xin, Wu, Yi, Yu, Xinjia, Zhang, Yan, Zhang, Lei, Wang, Di, Lyu, Shengfei, Xu, Hong, Wang, Xiaoqiao, Liu, Wei, Miao, Chunyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports
by: George, Sherine, et al.
Published: (2025)
by: George, Sherine, et al.
Published: (2025)
QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation
by: Lyu, Rita Qiuran, et al.
Published: (2026)
by: Lyu, Rita Qiuran, et al.
Published: (2026)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation
by: Ye, Hua, et al.
Published: (2026)
by: Ye, Hua, et al.
Published: (2026)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
by: Pan, Yating, et al.
Published: (2026)
by: Pan, Yating, et al.
Published: (2026)
Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
by: Severin, Nikita, et al.
Published: (2026)
by: Severin, Nikita, et al.
Published: (2026)
Neuromem: A Granular Decomposition of the Streaming Lifecycle in External Memory for LLMs
by: Zhang, Ruicheng, et al.
Published: (2026)
by: Zhang, Ruicheng, et al.
Published: (2026)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
by: Eyasir, Abrar, et al.
Published: (2026)
by: Eyasir, Abrar, et al.
Published: (2026)
Algorithmic Trust and Compliance: Benchmarking Brand Notability for UK iGaming Entities in Generative Search Engines
by: Oruesagasti, Julen
Published: (2026)
by: Oruesagasti, Julen
Published: (2026)
Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems
by: Cirillo, Stefano, et al.
Published: (2026)
by: Cirillo, Stefano, et al.
Published: (2026)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
by: Ahmed, Md Shamim, et al.
Published: (2026)
by: Ahmed, Md Shamim, et al.
Published: (2026)
AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation
by: Yao, Zhihui, et al.
Published: (2026)
by: Yao, Zhihui, et al.
Published: (2026)
When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
by: Shao, Jiaqi, et al.
Published: (2026)
by: Shao, Jiaqi, et al.
Published: (2026)
Graph-GRPO: Dependency-Aware Credit Assignment for Generative E-commerce Search Relevance
by: Che, Jiarui, et al.
Published: (2026)
by: Che, Jiarui, et al.
Published: (2026)
Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study
by: Tummalapenta, Ravi Kumar, et al.
Published: (2026)
by: Tummalapenta, Ravi Kumar, et al.
Published: (2026)
Leveraging LLMs to Create Content Corpora for Niche Domains
by: Zhang, Franklin, et al.
Published: (2025)
by: Zhang, Franklin, et al.
Published: (2025)
Augmented Relevance Datasets with Fine-Tuned Small LLMs
by: Fitte-Rey, Quentin, et al.
Published: (2025)
by: Fitte-Rey, Quentin, et al.
Published: (2025)
The Case for Intent-Based Query Rewriting
by: Nicolai, Gianna Lisa, et al.
Published: (2025)
by: Nicolai, Gianna Lisa, et al.
Published: (2025)
Motion Compensation for Real Time Ultrasound Scanning in Robotically Assisted Prostate Biopsy Procedures
by: Markulin, Matija, et al.
Published: (2026)
by: Markulin, Matija, et al.
Published: (2026)
Permanent Data Encoding (PDE): A Visual Language for Semantic Compression and Knowledge Preservation in 3-Character Units
by: Tsuyuki, Yoshiharu, et al.
Published: (2025)
by: Tsuyuki, Yoshiharu, et al.
Published: (2025)
Curated AI beats frontier LLMs at pharma asset discovery
by: Kidziński, Łukasz, et al.
Published: (2026)
by: Kidziński, Łukasz, et al.
Published: (2026)
Expanding Relevance Judgments for Medical Case-based Retrieval Task with Multimodal LLMs
by: Pires, Catarina, et al.
Published: (2025)
by: Pires, Catarina, et al.
Published: (2025)
From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
by: Kai, Zhang, et al.
Published: (2026)
by: Kai, Zhang, et al.
Published: (2026)
EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
by: Sun, Yuhong, et al.
Published: (2026)
by: Sun, Yuhong, et al.
Published: (2026)
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge Retrieval
by: Niu, Mengjia, et al.
Published: (2024)
by: Niu, Mengjia, et al.
Published: (2024)
Token-Oriented Object Notation vs JSON: A Benchmark of Plain and Constrained Decoding Generation
by: Matveev, Ivan
Published: (2026)
by: Matveev, Ivan
Published: (2026)
Harnessing multiple LLMs for Information Retrieval: A case study on Deep Learning methodologies in Biodiversity publications
by: Kommineni, Vamsi Krishna, et al.
Published: (2024)
by: Kommineni, Vamsi Krishna, et al.
Published: (2024)
Mind the Gap: Aligning Knowledge Bases with User Needs to Enhance Mental Health Retrieval
by: Chan, Amanda, et al.
Published: (2025)
by: Chan, Amanda, et al.
Published: (2025)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
by: Sakhovskiy, Andrey, et al.
Published: (2025)
by: Sakhovskiy, Andrey, et al.
Published: (2025)
A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
by: Mishra, Pranav Pushkar, et al.
Published: (2025)
by: Mishra, Pranav Pushkar, et al.
Published: (2025)
STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
by: Yang, Zhenye, et al.
Published: (2025)
by: Yang, Zhenye, et al.
Published: (2025)
RLM-on-KG: Heuristics First, LLMs When Needed: Adaptive Retrieval Control over Mention Graphs for Scattered Evidence
by: Volpini, Andrea, et al.
Published: (2026)
by: Volpini, Andrea, et al.
Published: (2026)
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
by: Böckling, Martin, et al.
Published: (2025)
by: Böckling, Martin, et al.
Published: (2025)
WisPaper: Your AI Scholar Search Engine
by: Ju, Li, et al.
Published: (2025)
by: Ju, Li, et al.
Published: (2025)
Comparison between the Structures of Word Co-occurrence and Word Similarity Networks for Ill-formed and Well-formed Texts in Taiwan Mandarin
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation
by: Zhang, Yongyue, et al.
Published: (2026)
by: Zhang, Yongyue, et al.
Published: (2026)
Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Memory as Metabolism: A Design for Companion Knowledge Systems
by: Miteski, Stefan
Published: (2026)
by: Miteski, Stefan
Published: (2026)
Flippi: End To End GenAI Assistant for E-Commerce
by: Rajasekar, Anand A., et al.
Published: (2025)
by: Rajasekar, Anand A., et al.
Published: (2025)
Similar Items
-
ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports
by: George, Sherine, et al.
Published: (2025) -
QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation
by: Lyu, Rita Qiuran, et al.
Published: (2026) -
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026) -
Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation
by: Ye, Hua, et al.
Published: (2026) -
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)