Guardado en:
| Autores principales: | Abdaljalil, Samir, Pallucchini, Filippo, Seveso, Andrea, Kurban, Hasan, Mercorio, Fabio, Serpedin, Erchin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2503.03032 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
por: Abdaljalil, Samir, et al.
Publicado: (2026)
por: Abdaljalil, Samir, et al.
Publicado: (2026)
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
por: Abdaljalil, Samir, et al.
Publicado: (2026)
por: Abdaljalil, Samir, et al.
Publicado: (2026)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
por: Polat, Can, et al.
Publicado: (2025)
por: Polat, Can, et al.
Publicado: (2025)
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2026)
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2026)
SCALAR: Quantifying Structural Hallucination, Consistency, and Reasoning Gaps in Materials Foundation Models
por: Polat, Can, et al.
Publicado: (2026)
por: Polat, Can, et al.
Publicado: (2026)
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
por: Mahmood, Alhasan, et al.
Publicado: (2026)
por: Mahmood, Alhasan, et al.
Publicado: (2026)
Designing Role Vectors to Improve LLM Inference Behaviour
por: Potertì, Daniele, et al.
Publicado: (2025)
por: Potertì, Daniele, et al.
Publicado: (2025)
QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions
por: Polat, Can, et al.
Publicado: (2025)
por: Polat, Can, et al.
Publicado: (2025)
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
por: Cambria, Erik, et al.
Publicado: (2024)
por: Cambria, Erik, et al.
Publicado: (2024)
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
por: Mercorio, Fabio, et al.
Publicado: (2024)
por: Mercorio, Fabio, et al.
Publicado: (2024)
Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework
por: Polat, Can, et al.
Publicado: (2025)
por: Polat, Can, et al.
Publicado: (2025)
xChemAgents: Agentic AI for Explainable Quantum Chemistry
por: Polat, Can, et al.
Publicado: (2025)
por: Polat, Can, et al.
Publicado: (2025)
C2NP: A Benchmark for Learning Scale-Dependent Geometric Invariances in 3D Materials Generation
por: Polat, Can, et al.
Publicado: (2026)
por: Polat, Can, et al.
Publicado: (2026)
Understanding the Capabilities of Molecular Graph Neural Networks in Materials Science Through Multimodal Learning and Physical Context Encoding
por: Polat, Can, et al.
Publicado: (2025)
por: Polat, Can, et al.
Publicado: (2025)
How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science
por: Polat, Can, et al.
Publicado: (2026)
por: Polat, Can, et al.
Publicado: (2026)
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
por: Khanbayov, Rasul, et al.
Publicado: (2026)
por: Khanbayov, Rasul, et al.
Publicado: (2026)
EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2025)
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
por: Hua, Zhenglin, et al.
Publicado: (2025)
por: Hua, Zhenglin, et al.
Publicado: (2025)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
por: Deng, Boyi, et al.
Publicado: (2025)
por: Deng, Boyi, et al.
Publicado: (2025)
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
por: Nandi, Palash, et al.
Publicado: (2024)
por: Nandi, Palash, et al.
Publicado: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
por: Assogba, Yannick, et al.
Publicado: (2026)
por: Assogba, Yannick, et al.
Publicado: (2026)
Joint Sensor Deployment and Physics-Informed Graph Transformer for Smart Grid Attack Detection
por: Elnour, Mariam, et al.
Publicado: (2026)
por: Elnour, Mariam, et al.
Publicado: (2026)
A Concise Review of Hallucinations in LLMs and their Mitigation
por: Pulkundwar, Parth, et al.
Publicado: (2025)
por: Pulkundwar, Parth, et al.
Publicado: (2025)
Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning
por: Liao, Yiming, et al.
Publicado: (2026)
por: Liao, Yiming, et al.
Publicado: (2026)
No One Size Fits All: QueryBandits for Hallucination Mitigation
por: Cho, Nicole, et al.
Publicado: (2026)
por: Cho, Nicole, et al.
Publicado: (2026)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
por: Goyal, Agam, et al.
Publicado: (2025)
por: Goyal, Agam, et al.
Publicado: (2025)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
por: Härle, Ruben, et al.
Publicado: (2024)
por: Härle, Ruben, et al.
Publicado: (2024)
Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
por: Ding, Hanxing, et al.
Publicado: (2024)
por: Ding, Hanxing, et al.
Publicado: (2024)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
por: Cho, Nicole, et al.
Publicado: (2025)
por: Cho, Nicole, et al.
Publicado: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
por: Saiyed, Ahson, et al.
Publicado: (2026)
por: Saiyed, Ahson, et al.
Publicado: (2026)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
por: Wu, Xuansheng, et al.
Publicado: (2025)
por: Wu, Xuansheng, et al.
Publicado: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
por: Sun, Zhongxiang, et al.
Publicado: (2026)
por: Sun, Zhongxiang, et al.
Publicado: (2026)
Natural Language Querying System Through Entity Enrichment
por: Amavi, Joshua, et al.
Publicado: (2024)
por: Amavi, Joshua, et al.
Publicado: (2024)
Mitigating Object Hallucination via Robust Local Perception Search
por: Gao, Zixian, et al.
Publicado: (2025)
por: Gao, Zixian, et al.
Publicado: (2025)
Ejemplares similares
-
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
por: Abdaljalil, Samir, et al.
Publicado: (2025) -
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
por: Abdaljalil, Samir, et al.
Publicado: (2026) -
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
por: Abdaljalil, Samir, et al.
Publicado: (2025) -
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
por: Abdaljalil, Samir, et al.
Publicado: (2026) -
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
por: Abdaljalil, Samir, et al.
Publicado: (2025)