Salvato in:
| Autore principale: | Advani, Laksh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2601.00513 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
di: Advani, Laksh
Pubblicazione: (2026)
di: Advani, Laksh
Pubblicazione: (2026)
Clever Materials: When Models Identify Good Materials for the Wrong Reasons
di: Jablonka, Kevin Maik
Pubblicazione: (2026)
di: Jablonka, Kevin Maik
Pubblicazione: (2026)
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
di: Patel, Laksh, et al.
Pubblicazione: (2025)
di: Patel, Laksh, et al.
Pubblicazione: (2025)
Wrong Model, Right Uncertainty: Spatial Associations for Discrete Data with Misspecification
di: Burt, David R., et al.
Pubblicazione: (2025)
di: Burt, David R., et al.
Pubblicazione: (2025)
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)
di: Veličković, Petar, et al.
Pubblicazione: (2026)
When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
di: Guo, Zhixiang, et al.
Pubblicazione: (2026)
di: Guo, Zhixiang, et al.
Pubblicazione: (2026)
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
di: Son, Seongho, et al.
Pubblicazione: (2024)
di: Son, Seongho, et al.
Pubblicazione: (2024)
Stable but Wrong: When More Data Degrades Scientific Conclusions
di: Zhang, Zhipeng, et al.
Pubblicazione: (2026)
di: Zhang, Zhipeng, et al.
Pubblicazione: (2026)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
When PINNs Go Wrong: Pseudo-Time Stepping Against Spurious Solutions
di: Wang, Sifan, et al.
Pubblicazione: (2026)
di: Wang, Sifan, et al.
Pubblicazione: (2026)
[Experiments & Analysis] Evaluating the Feasibility of Sampling-Based Techniques for Training Multilayer Perceptrons
di: Ebrahimi, Sana, et al.
Pubblicazione: (2023)
di: Ebrahimi, Sana, et al.
Pubblicazione: (2023)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents
di: Long, Carol Xuan
Pubblicazione: (2026)
di: Long, Carol Xuan
Pubblicazione: (2026)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
di: Singhi, Nishad, et al.
Pubblicazione: (2025)
di: Singhi, Nishad, et al.
Pubblicazione: (2025)
Thinking Wrong in Silence: Backdoor Attacks on Continuous Latent Reasoning
di: Parekh, Swapnil
Pubblicazione: (2026)
di: Parekh, Swapnil
Pubblicazione: (2026)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
di: Xiaohu, Xie, et al.
Pubblicazione: (2026)
di: Xiaohu, Xie, et al.
Pubblicazione: (2026)
When Learning Rates Go Wrong: Early Structural Signals in PPO Actor-Critic
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
di: Baum, Kevin, et al.
Pubblicazione: (2024)
di: Baum, Kevin, et al.
Pubblicazione: (2024)
Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
di: Andolfi, Luca, et al.
Pubblicazione: (2025)
di: Andolfi, Luca, et al.
Pubblicazione: (2025)
Towards Trustworthy GUI Agents: A Survey
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
di: Cai, Siyang, et al.
Pubblicazione: (2026)
di: Cai, Siyang, et al.
Pubblicazione: (2026)
Decidable By Construction: Design-Time Verification for Trustworthy AI
di: Haynes, Houston
Pubblicazione: (2026)
di: Haynes, Houston
Pubblicazione: (2026)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
di: Garcia, Gabriel
Pubblicazione: (2026)
di: Garcia, Gabriel
Pubblicazione: (2026)
All AI Models are Wrong, but Some are Optimal
di: Anand, Akhil S, et al.
Pubblicazione: (2025)
di: Anand, Akhil S, et al.
Pubblicazione: (2025)
LLMs as Assessors: Right for the Right Reason?
di: Saha, Sourav, et al.
Pubblicazione: (2026)
di: Saha, Sourav, et al.
Pubblicazione: (2026)
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
di: Nguyen, Thong Q., et al.
Pubblicazione: (2025)
di: Nguyen, Thong Q., et al.
Pubblicazione: (2025)
What is Wrong with Perplexity for Long-context Language Modeling?
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
Trustworthy Prediction with Gaussian Process Knowledge Scores
di: Butler, Kurt, et al.
Pubblicazione: (2025)
di: Butler, Kurt, et al.
Pubblicazione: (2025)
Verification and Validation for Trustworthy Scientific Machine Learning
di: Jakeman, John D., et al.
Pubblicazione: (2025)
di: Jakeman, John D., et al.
Pubblicazione: (2025)
Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision
di: Ning, Kanghui, et al.
Pubblicazione: (2025)
di: Ning, Kanghui, et al.
Pubblicazione: (2025)
Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
di: Martinez, John Ray B.
Pubblicazione: (2026)
di: Martinez, John Ray B.
Pubblicazione: (2026)
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
di: Wang, Hao, et al.
Pubblicazione: (2023)
di: Wang, Hao, et al.
Pubblicazione: (2023)
ManifoldMind: Dynamic Hyperbolic Reasoning for Trustworthy Recommendations
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
Step-by-Step Diffusion: An Elementary Tutorial
di: Nakkiran, Preetum, et al.
Pubblicazione: (2024)
di: Nakkiran, Preetum, et al.
Pubblicazione: (2024)
Out-of-Distribution Detection Methods Answer the Wrong Questions
di: Li, Yucen Lily, et al.
Pubblicazione: (2025)
di: Li, Yucen Lily, et al.
Pubblicazione: (2025)
Your Assumed DAG is Wrong and Here's How To Deal With It
di: Padh, Kirtan, et al.
Pubblicazione: (2025)
di: Padh, Kirtan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
di: Advani, Laksh
Pubblicazione: (2026) -
Clever Materials: When Models Identify Good Materials for the Wrong Reasons
di: Jablonka, Kevin Maik
Pubblicazione: (2026) -
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
di: Patel, Laksh, et al.
Pubblicazione: (2025) -
Wrong Model, Right Uncertainty: Spatial Associations for Discrete Data with Misspecification
di: Burt, David R., et al.
Pubblicazione: (2025) -
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)