When Counterfactual Reasoning Fails: Chaos and Real-World Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | Aalaila, Yahya, Großmann, Gerrit, Mukherjee, Sumantrak, Wahl, Jonas, Vollmer, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Spatiotemporal Point Processes: Trends and Challenges
by: Mukherjee, Sumantrak, et al.
Published: (2025)
by: Mukherjee, Sumantrak, et al.
Published: (2025)
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
by: Mukherjee, Sumantrak, et al.
Published: (2025)
by: Mukherjee, Sumantrak, et al.
Published: (2025)
Guiding Exploration in Reinforcement Learning Through LLM-Augmented Observations
by: Jain, Vaibhav, et al.
Published: (2025)
by: Jain, Vaibhav, et al.
Published: (2025)
Auto-encoding Molecules: Graph-Matching Capabilities Matter
by: Cunow, Magnus, et al.
Published: (2025)
by: Cunow, Magnus, et al.
Published: (2025)
Exploring Molecule Generation Using Latent Space Graph Diffusion
by: Pombala, Prashanth, et al.
Published: (2025)
by: Pombala, Prashanth, et al.
Published: (2025)
Enhancing GNNs with Architecture-Agnostic Graph Transformations: A Systematic Analysis
by: Li, Zhifei, et al.
Published: (2024)
by: Li, Zhifei, et al.
Published: (2024)
Graph Agnostic Causal Bayesian Optimisation
by: Mukherjee, Sumantrak, et al.
Published: (2024)
by: Mukherjee, Sumantrak, et al.
Published: (2024)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
by: Kwon, Yongchan, et al.
Published: (2025)
by: Kwon, Yongchan, et al.
Published: (2025)
When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
by: Park, Min-Yeong, et al.
Published: (2025)
by: Park, Min-Yeong, et al.
Published: (2025)
Critical Windows of Complexity Control: When Transformers Decide to Reason or Memorize
by: Ali, Sarwan
Published: (2026)
by: Ali, Sarwan
Published: (2026)
Frequency Matters: When Time Series Foundation Models Fail Under Spectral Shift
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift
by: Vogt-Lowell, Kevin, et al.
Published: (2026)
by: Vogt-Lowell, Kevin, et al.
Published: (2026)
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
by: Erol, Melihcan, et al.
Published: (2026)
by: Erol, Melihcan, et al.
Published: (2026)
MEDAKA: Construction of Biomedical Knowledge Graphs Using Large Language Models
by: Sengupta, Asmita, et al.
Published: (2025)
by: Sengupta, Asmita, et al.
Published: (2025)
Improving Counterfactual Truthfulness for Molecular Property Prediction through Uncertainty Quantification
by: Teufel, Jonas, et al.
Published: (2025)
by: Teufel, Jonas, et al.
Published: (2025)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
by: Wang, Yongxin, et al.
Published: (2025)
by: Wang, Yongxin, et al.
Published: (2025)
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
Properties that allow or prohibit transferability of adversarial attacks among quantized networks
by: Shrestha, Abhishek, et al.
Published: (2024)
by: Shrestha, Abhishek, et al.
Published: (2024)
When Mean CE Fails: Median CE Can Better Track Language Model Quality
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning
by: Jerge, Michael, et al.
Published: (2026)
by: Jerge, Michael, et al.
Published: (2026)
When Validation Fails: Cross-Institutional Blood Pressure Prediction and the Limits of Electronic Health Record-Based Models
by: Azam, Md Basit, et al.
Published: (2025)
by: Azam, Md Basit, et al.
Published: (2025)
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
by: Renda, Alana, et al.
Published: (2025)
by: Renda, Alana, et al.
Published: (2025)
Benchmarking World-Model Learning with Environment-Level Queries
by: Warrier, Archana, et al.
Published: (2025)
by: Warrier, Archana, et al.
Published: (2025)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
by: Guo, Hanze, et al.
Published: (2025)
by: Guo, Hanze, et al.
Published: (2025)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning
by: Koka, Yousef, et al.
Published: (2025)
by: Koka, Yousef, et al.
Published: (2025)
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
by: Chandra, Abhranil, et al.
Published: (2025)
by: Chandra, Abhranil, et al.
Published: (2025)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
by: Landesberg, Eddie
Published: (2026)
by: Landesberg, Eddie
Published: (2026)
RealAC: A Domain-Agnostic Framework for Realistic and Actionable Counterfactual Explanations
by: Arefeen, Asiful, et al.
Published: (2025)
by: Arefeen, Asiful, et al.
Published: (2025)
Current Agents Fail to Leverage World Model as Tool for Foresight
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Augmenting The Weather: A Hybrid Counterfactual-SMOTE Algorithm for Improving Crop Growth Prediction When Climate Changes
by: Temraz, Mohammed, et al.
Published: (2025)
by: Temraz, Mohammed, et al.
Published: (2025)
Similar Items
-
Neural Spatiotemporal Point Processes: Trends and Challenges
by: Mukherjee, Sumantrak, et al.
Published: (2025) -
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
by: Mukherjee, Sumantrak, et al.
Published: (2025) -
Guiding Exploration in Reinforcement Learning Through LLM-Augmented Observations
by: Jain, Vaibhav, et al.
Published: (2025) -
Auto-encoding Molecules: Graph-Matching Capabilities Matter
by: Cunow, Magnus, et al.
Published: (2025) -
Exploring Molecule Generation Using Latent Space Graph Diffusion
by: Pombala, Prashanth, et al.
Published: (2025)