Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Zihan, Datta, Rituparna, Hu, Mengxuan, Liu, Shunshun, Zhang, Aiying, Balachandran, Prasanna, Li, Sheng, Vullikanti, Anil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
by: Guan, Zihan, et al.
Published: (2025)
by: Guan, Zihan, et al.
Published: (2025)
UFID: A Unified Framework for Input-level Backdoor Detection on Diffusion Models
by: Guan, Zihan, et al.
Published: (2024)
by: Guan, Zihan, et al.
Published: (2024)
Agentic Framework for Epidemiological Modeling
by: Datta, Rituparna, et al.
Published: (2026)
by: Datta, Rituparna, et al.
Published: (2026)
Large Language Models Lack Temporal Awareness of Medical Knowledge
by: Guan, Zihan, et al.
Published: (2026)
by: Guan, Zihan, et al.
Published: (2026)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
CALYPSO: Forecasting and Analyzing MRSA Infection Patterns with Community and Healthcare Transmission Dynamics
by: Datta, Rituparna, et al.
Published: (2025)
by: Datta, Rituparna, et al.
Published: (2025)
Large Language Models for Causal Discovery: Current Landscape and Future Directions
by: Wan, Guangya, et al.
Published: (2024)
by: Wan, Guangya, et al.
Published: (2024)
BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing
by: Guo, Dongliang, et al.
Published: (2025)
by: Guo, Dongliang, et al.
Published: (2025)
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
by: Hu, Mengxuan, et al.
Published: (2024)
by: Hu, Mengxuan, et al.
Published: (2024)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
by: Shou, Yitong, et al.
Published: (2026)
by: Shou, Yitong, et al.
Published: (2026)
Enhanced Diagnostic Performance via Large-Resolution Inference Optimization for Pathology Foundation Models
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2024)
by: Li, Shuyue Stella, et al.
Published: (2024)
Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling
by: Datta, Rituparna, et al.
Published: (2025)
by: Datta, Rituparna, et al.
Published: (2025)
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing
by: Guo, Dongliang, et al.
Published: (2024)
by: Guo, Dongliang, et al.
Published: (2024)
Are LLMs Ready to Replace Bangla Annotators?
by: Hasan, Md. Najib, et al.
Published: (2026)
by: Hasan, Md. Najib, et al.
Published: (2026)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Differentially private exact recovery for stochastic block models
by: Nguyen, Dung, et al.
Published: (2024)
by: Nguyen, Dung, et al.
Published: (2024)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
by: Sun, Moran, et al.
Published: (2026)
by: Sun, Moran, et al.
Published: (2026)
Differentially Private Densest Subgraph Detection
by: Nguyen, Dung, et al.
Published: (2021)
by: Nguyen, Dung, et al.
Published: (2021)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
by: Qi, Yunjia, et al.
Published: (2025)
by: Qi, Yunjia, et al.
Published: (2025)
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025)
by: Jiang, Bohan, et al.
Published: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Orchard: An Open-Source Agentic Modeling Framework
by: Peng, Baolin, et al.
Published: (2026)
by: Peng, Baolin, et al.
Published: (2026)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
by: Wu, Junde, et al.
Published: (2025)
by: Wu, Junde, et al.
Published: (2025)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios
by: Zhao, Bingxi, et al.
Published: (2025)
by: Zhao, Bingxi, et al.
Published: (2025)
GEM: A Gym for Agentic LLMs
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
The Evolution of RWKV: Advancements in Efficient Language Modeling
by: Datta, Akul
Published: (2024)
by: Datta, Akul
Published: (2024)
Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
by: Li, Yihan, et al.
Published: (2025)
by: Li, Yihan, et al.
Published: (2025)
LLMs as Agentic Cooperative Players in Multiplayer UNO
by: Matinez, Yago Romano, et al.
Published: (2025)
by: Matinez, Yago Romano, et al.
Published: (2025)
Similar Items
-
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
by: Guan, Zihan, et al.
Published: (2025) -
UFID: A Unified Framework for Input-level Backdoor Detection on Diffusion Models
by: Guan, Zihan, et al.
Published: (2024) -
Agentic Framework for Epidemiological Modeling
by: Datta, Rituparna, et al.
Published: (2026) -
Large Language Models Lack Temporal Awareness of Medical Knowledge
by: Guan, Zihan, et al.
Published: (2026) -
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)