Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Srivastava, Ashutosh, Nagalapatti, Lokesh, Jajoo, Gautam, Vashishtha, Aniket, Krishnamurthy, Parameswari, Sharma, Amit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Root Cause Diagnosis using In-Distribution Interventions
by: Nagalapatti, Lokesh, et al.
Published: (2025)
by: Nagalapatti, Lokesh, et al.
Published: (2025)
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
by: Vashishtha, Aniket, et al.
Published: (2025)
by: Vashishtha, Aniket, et al.
Published: (2025)
Leveraging a Simulator for Learning Causal Representations from Post-Treatment Covariates for CATE
by: Nagalapatti, Lokesh, et al.
Published: (2025)
by: Nagalapatti, Lokesh, et al.
Published: (2025)
From Search To Sampling: Generative Models For Robust Algorithmic Recourse
by: Garg, Prateek, et al.
Published: (2025)
by: Garg, Prateek, et al.
Published: (2025)
PairNet: Training with Observed Pairs to Estimate Individual Treatment Effect
by: Nagalapatti, Lokesh, et al.
Published: (2024)
by: Nagalapatti, Lokesh, et al.
Published: (2024)
Gradient Coreset for Federated Learning
by: Sivasubramanian, Durga, et al.
Published: (2024)
by: Sivasubramanian, Durga, et al.
Published: (2024)
Continuous Treatment Effect Estimation Using Gradient Interpolation and Kernel Smoothing
by: Nagalapatti, Lokesh, et al.
Published: (2024)
by: Nagalapatti, Lokesh, et al.
Published: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
On the Internal Semantics of Time-Series Foundation Models
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
Task Facet Learning: A Structured Approach to Prompt Optimization
by: Juneja, Gurusha, et al.
Published: (2024)
by: Juneja, Gurusha, et al.
Published: (2024)
Tab-Shapley: Identifying Top-k Tabular Data Quality Insights
by: Padala, Manisha, et al.
Published: (2025)
by: Padala, Manisha, et al.
Published: (2025)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
by: Singh, Sanjeet, et al.
Published: (2024)
by: Singh, Sanjeet, et al.
Published: (2024)
MASCA: LLM based-Multi Agents System for Credit Assessment
by: Jajoo, Gautam, et al.
Published: (2025)
by: Jajoo, Gautam, et al.
Published: (2025)
Incremental Multi-Scene Modeling via Continual Neural Graphics Primitives
by: Singh, Prajwal, et al.
Published: (2024)
by: Singh, Prajwal, et al.
Published: (2024)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
by: Zhou, Yu, et al.
Published: (2024)
by: Zhou, Yu, et al.
Published: (2024)
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
NICE: To Optimize In-Context Examples or Not?
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
LLMs are Bayesian, In Expectation, Not in Realization
by: Chlon, Leon, et al.
Published: (2025)
by: Chlon, Leon, et al.
Published: (2025)
Improving Generative Methods for Causal Evaluation via Simulation-Based Inference
by: Amaranath, Pracheta, et al.
Published: (2025)
by: Amaranath, Pracheta, et al.
Published: (2025)
The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
by: DeLeeuw, Caleb, et al.
Published: (2025)
by: DeLeeuw, Caleb, et al.
Published: (2025)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
by: Vashishtha, Aniket, et al.
Published: (2023)
by: Vashishtha, Aniket, et al.
Published: (2023)
Mamba Outpaces Reformer in Stock Prediction with Sentiments from Top Ten LLMs
by: Kadiyala, Lokesh Antony, et al.
Published: (2025)
by: Kadiyala, Lokesh Antony, et al.
Published: (2025)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
by: Hsu, Sheryl, et al.
Published: (2024)
by: Hsu, Sheryl, et al.
Published: (2024)
Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models
by: Yamaoka, Kanta, et al.
Published: (2026)
by: Yamaoka, Kanta, et al.
Published: (2026)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
by: Krishnamurthy, Vikram
Published: (2026)
by: Krishnamurthy, Vikram
Published: (2026)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
by: Sharma, Amit
Published: (2025)
by: Sharma, Amit
Published: (2025)
CausalPlayground: Addressing Data-Generation Requirements in Cutting-Edge Causality Research
by: Sauter, Andreas W M, et al.
Published: (2024)
by: Sauter, Andreas W M, et al.
Published: (2024)
ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)
by: Singhal, Kartik, et al.
Published: (2024)
by: Singhal, Kartik, et al.
Published: (2024)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models
by: Ranjan, Ashutosh, et al.
Published: (2026)
by: Ranjan, Ashutosh, et al.
Published: (2026)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
Creating a Causally Grounded Rating Method for Assessing the Robustness of AI Models for Time-Series Forecasting
by: Lakkaraju, Kausik, et al.
Published: (2025)
by: Lakkaraju, Kausik, et al.
Published: (2025)
Yantra AI -- An intelligence platform which interacts with manufacturing operations
by: Krishnamurthy, Varshini
Published: (2025)
by: Krishnamurthy, Varshini
Published: (2025)
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and Challenges
by: Quan, Pengrui, et al.
Published: (2025)
by: Quan, Pengrui, et al.
Published: (2025)
Benchmarking Reliability of Deep Learning Models for Pathological Gait Classification
by: Jaiswal, Abhishek, et al.
Published: (2024)
by: Jaiswal, Abhishek, et al.
Published: (2024)
Improving Self Consistency in LLMs through Probabilistic Tokenization
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
by: Pang, Jiacheng, et al.
Published: (2026)
by: Pang, Jiacheng, et al.
Published: (2026)
Position: Understanding LLMs Requires More Than Statistical Generalization
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Similar Items
-
Robust Root Cause Diagnosis using In-Distribution Interventions
by: Nagalapatti, Lokesh, et al.
Published: (2025) -
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
by: Vashishtha, Aniket, et al.
Published: (2025) -
Leveraging a Simulator for Learning Causal Representations from Post-Treatment Covariates for CATE
by: Nagalapatti, Lokesh, et al.
Published: (2025) -
From Search To Sampling: Generative Models For Robust Algorithmic Recourse
by: Garg, Prateek, et al.
Published: (2025) -
PairNet: Training with Observed Pairs to Estimate Individual Treatment Effect
by: Nagalapatti, Lokesh, et al.
Published: (2024)