ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Basu, Abhinaba, Chakraborty, Pavan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
by: Lin, Shuhang, et al.
Published: (2026)
by: Lin, Shuhang, et al.
Published: (2026)
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
by: de Zuazo, Xabier, et al.
Published: (2025)
by: de Zuazo, Xabier, et al.
Published: (2025)
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
by: Abtahi, Farhad, et al.
Published: (2026)
by: Abtahi, Farhad, et al.
Published: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Latent-Autoregressive GP-VAE Language Model
by: Ruffenach, Yves
Published: (2025)
by: Ruffenach, Yves
Published: (2025)
Conformal Prediction Sets for Next-Token Prediction in Large Language Models: Balancing Coverage Guarantees with Set Efficiency
by: Kotla, Yoshith Roy, et al.
Published: (2025)
by: Kotla, Yoshith Roy, et al.
Published: (2025)
Classifier Calibration at Scale: An Empirical Study of Model-Agnostic Post-Hoc Methods
by: Manokhin, Valery, et al.
Published: (2026)
by: Manokhin, Valery, et al.
Published: (2026)
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
by: Badshah, Sher, et al.
Published: (2026)
by: Badshah, Sher, et al.
Published: (2026)
ML-driven detection and reduction of ballast information in multi-modal datasets
by: Solovko, Yaroslav
Published: (2026)
by: Solovko, Yaroslav
Published: (2026)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
by: Chang, Sidi, et al.
Published: (2026)
by: Chang, Sidi, et al.
Published: (2026)
A Hybrid Framework for Real-Time Data Drift and Anomaly Identification Using Hierarchical Temporal Memory and Statistical Tests
by: Bandyopadhyay, Subhadip, et al.
Published: (2025)
by: Bandyopadhyay, Subhadip, et al.
Published: (2025)
Advancements in Machine Learning and Deep Learning for Early Detection and Management of Mental Health Disorder
by: Kannan, Kamala Devi, et al.
Published: (2024)
by: Kannan, Kamala Devi, et al.
Published: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
GIM: Evaluating models via tasks that integrate multiple cognitive domains
by: Patel, Rohit, et al.
Published: (2026)
by: Patel, Rohit, et al.
Published: (2026)
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
by: Du, Jin, et al.
Published: (2025)
by: Du, Jin, et al.
Published: (2025)
Closed-Form Beta Distribution Estimation from Sparse Statistics with Random Forest Implicit Regularization
by: Landers, Jonathan R.
Published: (2025)
by: Landers, Jonathan R.
Published: (2025)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
Claim Automation using Large Language Model
by: Mo, Zhengda, et al.
Published: (2026)
by: Mo, Zhengda, et al.
Published: (2026)
Conservative Decisions with Risk Scores
by: Wei, Yishu, et al.
Published: (2025)
by: Wei, Yishu, et al.
Published: (2025)
Optimised Support Vector Regression for California Housing Price Prediction: The Critical Role of Feature Engineering and Hyperparameter Tuning
by: Adutwum, Emmanuel
Published: (2026)
by: Adutwum, Emmanuel
Published: (2026)
Augmented Functional Random Forests: Classifier Construction and Unbiased Functional Principal Components Importance through Ad-Hoc Conditional Permutations
by: Maturo, Fabrizio, et al.
Published: (2024)
by: Maturo, Fabrizio, et al.
Published: (2024)
When Are Two RLHF Objectives the Same?
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability
by: Allen, Ryan, et al.
Published: (2026)
by: Allen, Ryan, et al.
Published: (2026)
Why Agent Caching Fails and How to Fix It: Structured Intent Canonicalization with Few-Shot Learning
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Text Clustering with Large Language Model Embeddings
by: Petukhova, Alina, et al.
Published: (2024)
by: Petukhova, Alina, et al.
Published: (2024)
J6: Jacobian-Driven Role Attribution for Multi-Objective Prompt Optimization in LLMs
by: Wu, Yao
Published: (2025)
by: Wu, Yao
Published: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
by: Wan, Shu, et al.
Published: (2026)
by: Wan, Shu, et al.
Published: (2026)
Federated Learning and Class Imbalances
by: Zhu, Siqi, et al.
Published: (2026)
by: Zhu, Siqi, et al.
Published: (2026)
SplitWise Regression: Stepwise Modeling with Adaptive Dummy Encoding
by: Kurbucz, Marcell T., et al.
Published: (2025)
by: Kurbucz, Marcell T., et al.
Published: (2025)
Asymptotic Consistency and Generalization in Hybrid Models of Regularized Selection and Nonlinear Learning
by: Galvão, Luciano Ribeiro, et al.
Published: (2025)
by: Galvão, Luciano Ribeiro, et al.
Published: (2025)
Cross-Domain Uncertainty Quantification for Selective Prediction: A Comprehensive Bound Ablation with Transfer-Informed Betting
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
The Trilemma of Truth in Large Language Models
by: Savcisens, Germans, et al.
Published: (2025)
by: Savcisens, Germans, et al.
Published: (2025)
Similar Items
-
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
by: Lin, Shuhang, et al.
Published: (2026) -
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
by: de Zuazo, Xabier, et al.
Published: (2025) -
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
by: Abtahi, Farhad, et al.
Published: (2026) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026) -
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
by: Gao, Yuxuan, et al.
Published: (2026)