Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdaljalil, Samir, Serpedin, Erchin, Qaraqe, Khalid, Kurban, Hasan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
SCALAR: Quantifying Structural Hallucination, Consistency, and Reasoning Gaps in Materials Foundation Models
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
Understanding the Capabilities of Molecular Graph Neural Networks in Materials Science Through Multimodal Learning and Physical Context Encoding
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
xChemAgents: Agentic AI for Explainable Quantum Chemistry
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
C2NP: A Benchmark for Learning Scale-Dependent Geometric Invariances in 3D Materials Generation
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
Syntactic Control of Language Models by Posterior Inference
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
PhysicsEval: Inference-Time Techniques to Improve the Reasoning Proficiency of Large Language Models on Physics Problems
von: Siddique, Oshayer, et al.
Veröffentlicht: (2025)
von: Siddique, Oshayer, et al.
Veröffentlicht: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
von: Mishra, Shubhra, et al.
Veröffentlicht: (2024)
von: Mishra, Shubhra, et al.
Veröffentlicht: (2024)
A Survey on Large Language Models for Mathematical Reasoning
von: Wang, Peng-Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Peng-Yuan, et al.
Veröffentlicht: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
von: Khanbayov, Rasul, et al.
Veröffentlicht: (2026)
von: Khanbayov, Rasul, et al.
Veröffentlicht: (2026)
Examining False Positives under Inference Scaling for Mathematical Reasoning
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
Large Language and Reasoning Models are Shallow Disjunctive Reasoners
von: Khalid, Irtaza, et al.
Veröffentlicht: (2025)
von: Khalid, Irtaza, et al.
Veröffentlicht: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models
von: Hossain, Jawad, et al.
Veröffentlicht: (2026)
von: Hossain, Jawad, et al.
Veröffentlicht: (2026)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Joint Sensor Deployment and Physics-Informed Graph Transformer for Smart Grid Attack Detection
von: Elnour, Mariam, et al.
Veröffentlicht: (2026)
von: Elnour, Mariam, et al.
Veröffentlicht: (2026)
On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
von: Ren, Siyu, et al.
Veröffentlicht: (2024)
von: Ren, Siyu, et al.
Veröffentlicht: (2024)
AuditWen:An Open-Source Large Language Model for Audit
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
von: Amjad, Husnain, et al.
Veröffentlicht: (2026)
von: Amjad, Husnain, et al.
Veröffentlicht: (2026)
GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models
von: Shi, Nannan, et al.
Veröffentlicht: (2025)
von: Shi, Nannan, et al.
Veröffentlicht: (2025)
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2025)
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2025)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026) -
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)