ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Cekinmez, Jasin, Ghahroodi, Omid, Chandle, Saad Fowad, Gupta, Dhiman, Asgari, Ehsaneddin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025)
by: Shankarampeta, Abhilash, et al.
Published: (2025)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
by: Pandya, Pranshu, et al.
Published: (2024)
by: Pandya, Pranshu, et al.
Published: (2024)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
by: Gupta, Vatsal, et al.
Published: (2023)
by: Gupta, Vatsal, et al.
Published: (2023)
Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG
by: Khadilkar, Harshad, et al.
Published: (2025)
by: Khadilkar, Harshad, et al.
Published: (2025)
TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition
by: Nahid, Md Mahadi Hasan, et al.
Published: (2024)
by: Nahid, Md Mahadi Hasan, et al.
Published: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
by: Chen, Hung-Ting, et al.
Published: (2024)
by: Chen, Hung-Ting, et al.
Published: (2024)
Redefining Retrieval Evaluation in the Era of LLMs
by: Trappolini, Giovanni, et al.
Published: (2025)
by: Trappolini, Giovanni, et al.
Published: (2025)
DistRAG: Towards Distance-Based Spatial Reasoning in LLMs
by: Schneider, Nicole R, et al.
Published: (2025)
by: Schneider, Nicole R, et al.
Published: (2025)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
by: Chen, Jianlyu, et al.
Published: (2025)
by: Chen, Jianlyu, et al.
Published: (2025)
Evaluating LLMs for Gender Disparities in Notable Persons
by: Rhue, Lauren, et al.
Published: (2024)
by: Rhue, Lauren, et al.
Published: (2024)
HyperG: Hypergraph-Enhanced LLMs for Structured Knowledge
by: Huang, Sirui, et al.
Published: (2025)
by: Huang, Sirui, et al.
Published: (2025)
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana
by: Filice, Simone, et al.
Published: (2025)
by: Filice, Simone, et al.
Published: (2025)
Evaluating the Performance of LLMs on Technical Language Processing tasks
by: Kernycky, Andrew, et al.
Published: (2024)
by: Kernycky, Andrew, et al.
Published: (2024)
Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering
by: Xu, Tao
Published: (2026)
by: Xu, Tao
Published: (2026)
Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
by: Liu, Qijiong, et al.
Published: (2025)
by: Liu, Qijiong, et al.
Published: (2025)
Text2Cypher Across Languages: Evaluating and Finetuning LLMs
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
by: Zhang, Longxiang, et al.
Published: (2026)
by: Zhang, Longxiang, et al.
Published: (2026)
Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
by: Qian, Hongjin, et al.
Published: (2025)
by: Qian, Hongjin, et al.
Published: (2025)
Knowledge in Triples for LLMs: Enhancing Table QA Accuracy with Semantic Extraction
by: Sholehrasa, Hossein, et al.
Published: (2024)
by: Sholehrasa, Hossein, et al.
Published: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
by: Martin, Alexander, et al.
Published: (2025)
by: Martin, Alexander, et al.
Published: (2025)
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
by: Bachyr, Omar El, et al.
Published: (2026)
by: Bachyr, Omar El, et al.
Published: (2026)
Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework
by: Buskila, Avi-ad Avraam
Published: (2026)
by: Buskila, Avi-ad Avraam
Published: (2026)
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
by: Guo, Minghao, et al.
Published: (2026)
by: Guo, Minghao, et al.
Published: (2026)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
by: Song, Tingyu, et al.
Published: (2026)
by: Song, Tingyu, et al.
Published: (2026)
MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models
by: Chen, Zhongpu, et al.
Published: (2025)
by: Chen, Zhongpu, et al.
Published: (2025)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
SRR-Judge: Step-Level Rating and Refinement for Enhancing Search-Integrated Reasoning in Search Agents
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
LLMs as Assessors: Right for the Right Reason?
by: Saha, Sourav, et al.
Published: (2026)
by: Saha, Sourav, et al.
Published: (2026)
ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning
by: Zhu, Changtai, et al.
Published: (2025)
by: Zhu, Changtai, et al.
Published: (2025)
Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability
by: B, Gautam, et al.
Published: (2024)
by: B, Gautam, et al.
Published: (2024)
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
by: Ma, Yongqiang, et al.
Published: (2024)
by: Ma, Yongqiang, et al.
Published: (2024)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
Similar Items
-
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025) -
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024) -
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025) -
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026) -
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
by: Pandya, Pranshu, et al.
Published: (2024)