Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Anuyah, Sydney, Shajee-Mohan, Sneha, Chauhan, Ankit-Singh, Chakraborty, Sunandan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study of Causal Relation Extraction Transfer: Design and Data
by: Anuyah, Sydney, et al.
Published: (2025)
by: Anuyah, Sydney, et al.
Published: (2025)
Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs
by: Anuyah, Sydney, et al.
Published: (2026)
by: Anuyah, Sydney, et al.
Published: (2026)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
by: Vakilian, Vala, et al.
Published: (2025)
by: Vakilian, Vala, et al.
Published: (2025)
Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
by: Anuyah, Sydney, et al.
Published: (2025)
by: Anuyah, Sydney, et al.
Published: (2025)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Cross-Domain Content Generation with Domain-Specific Small Language Models
by: Maloo, Ankit, et al.
Published: (2024)
by: Maloo, Ankit, et al.
Published: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
by: Li, Keyu, et al.
Published: (2025)
by: Li, Keyu, et al.
Published: (2025)
CARE: A QLoRA-Fine Tuned Multi-Domain Chatbot With Fast Learning On Minimal Hardware
by: Dutta, Ankit, et al.
Published: (2025)
by: Dutta, Ankit, et al.
Published: (2025)
Multi-Agent Causal Discovery Using Large Language Models
by: Le, Hao Duong, et al.
Published: (2024)
by: Le, Hao Duong, et al.
Published: (2024)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
MMRAG: Multi-Mode Retrieval-Augmented Generation with Large Language Models for Biomedical In-Context Learning
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
Leveraging Multi-AI Agents for Cross-Domain Knowledge Discovery
by: Aryal, Shiva, et al.
Published: (2024)
by: Aryal, Shiva, et al.
Published: (2024)
Reduction of Supervision for Biomedical Knowledge Discovery
by: Theodoropoulos, Christos, et al.
Published: (2025)
by: Theodoropoulos, Christos, et al.
Published: (2025)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
by: Lee, Donggyu, et al.
Published: (2025)
by: Lee, Donggyu, et al.
Published: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
by: Chang, Chen-Chi, et al.
Published: (2024)
by: Chang, Chen-Chi, et al.
Published: (2024)
NormSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-Fly
by: Fung, Yi R., et al.
Published: (2022)
by: Fung, Yi R., et al.
Published: (2022)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
HiFACTMix: A Code-Mixed Benchmark and Graph-Aware Model for EvidenceBased Political Claim Verification in Hinglish
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering
by: Parekh, Jash Rajesh, et al.
Published: (2026)
by: Parekh, Jash Rajesh, et al.
Published: (2026)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
by: Wang, Shouju, et al.
Published: (2026)
by: Wang, Shouju, et al.
Published: (2026)
A Benchmark for Cross-Domain Argumentative Stance Classification on Social Media
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Can We Edit LLMs for Long-Tail Biomedical Knowledge?
by: Yi, Xinhao, et al.
Published: (2025)
by: Yi, Xinhao, et al.
Published: (2025)
AI-assisted Knowledge Discovery in Biomedical Literature to Support Decision-making in Precision Oncology
by: He, Ting, et al.
Published: (2024)
by: He, Ting, et al.
Published: (2024)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
Beyond Continuity: Challenges of Context Switching in Multi-Turn Dialogue with LLMs
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Large Language Models for Constrained-Based Causal Discovery
by: Cohrs, Kai-Hendrik, et al.
Published: (2024)
by: Cohrs, Kai-Hendrik, et al.
Published: (2024)
DSBC : Data Science task Benchmarking with Context engineering
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
DrKGC: Dynamic Subgraph Retrieval-Augmented LLMs for Knowledge Graph Completion across General and Biomedical Domains
by: Xiao, Yongkang, et al.
Published: (2025)
by: Xiao, Yongkang, et al.
Published: (2025)
Causal Understanding by LLMs: The Role of Uncertainty
by: Lithgow-Serrano, Oscar, et al.
Published: (2025)
by: Lithgow-Serrano, Oscar, et al.
Published: (2025)
LLMs Are Prone to Fallacies in Causal Inference
by: Joshi, Nitish, et al.
Published: (2024)
by: Joshi, Nitish, et al.
Published: (2024)
Similar Items
-
An Empirical Study of Causal Relation Extraction Transfer: Design and Data
by: Anuyah, Sydney, et al.
Published: (2025) -
Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs
by: Anuyah, Sydney, et al.
Published: (2026) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024) -
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025) -
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
by: Vakilian, Vala, et al.
Published: (2025)