Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Ganguly, Debargha, Singh, Vikash, Sankar, Sreehari, Zhang, Biyao, Zhang, Xuecen, Iyengar, Srinivasan, Han, Xiaotian, Sharma, Amit, Kalyanaraman, Shivkumar, Chaudhary, Vipin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning
by: Ganguly, Debargha, et al.
Published: (2024)
by: Ganguly, Debargha, et al.
Published: (2024)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
by: Zhang, Biyao, et al.
Published: (2025)
by: Zhang, Biyao, et al.
Published: (2025)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
Trust The Typical
by: Ganguly, Debargha, et al.
Published: (2026)
by: Ganguly, Debargha, et al.
Published: (2026)
LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection
by: Zhang, Xuecen, et al.
Published: (2026)
by: Zhang, Xuecen, et al.
Published: (2026)
An Agentic Approach to Automatic Creation of P&ID Diagrams from Natural Language Descriptions
by: Gowaikar, Shreeyash, et al.
Published: (2024)
by: Gowaikar, Shreeyash, et al.
Published: (2024)
Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification
by: Singh, Vikash, et al.
Published: (2026)
by: Singh, Vikash, et al.
Published: (2026)
Visual Concept Networks: A Graph-Based Approach to Detecting Anomalous Data in Deep Neural Networks
by: Ganguly, Debargha, et al.
Published: (2024)
by: Ganguly, Debargha, et al.
Published: (2024)
SPIRIT: Short-term Prediction of solar IRradIance for zero-shot Transfer learning using Foundation Models
by: Mishra, Aditya, et al.
Published: (2025)
by: Mishra, Aditya, et al.
Published: (2025)
$K^4$: Online Log Anomaly Detection Via Unsupervised Typicality Learning
by: Chen, Weicong, et al.
Published: (2025)
by: Chen, Weicong, et al.
Published: (2025)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Reliability-Gated Source Anchoring for Continual Test-Time Adaptation
by: Singh, Vikash, et al.
Published: (2026)
by: Singh, Vikash, et al.
Published: (2026)
Forte : Finding Outliers with Representation Typicality Estimation
by: Ganguly, Debargha, et al.
Published: (2024)
by: Ganguly, Debargha, et al.
Published: (2024)
CausalGuard: Conformal Inference under Graph Uncertainty
by: Singh, Vikash, et al.
Published: (2026)
by: Singh, Vikash, et al.
Published: (2026)
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
by: Wang, Shouren, et al.
Published: (2026)
by: Wang, Shouren, et al.
Published: (2026)
Novel adaptation of video segmentation to 3D MRI: efficient zero-shot knee segmentation with SAM2
by: Yu, Andrew Seohwan, et al.
Published: (2024)
by: Yu, Andrew Seohwan, et al.
Published: (2024)
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Context Determines Optimal Architecture in Materials Segmentation
by: Lu, Mingjian, et al.
Published: (2026)
by: Lu, Mingjian, et al.
Published: (2026)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
by: Zafar, Osama, et al.
Published: (2026)
by: Zafar, Osama, et al.
Published: (2026)
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation
by: Wang, Nengbo, et al.
Published: (2025)
by: Wang, Nengbo, et al.
Published: (2025)
EnCortex: A General, Extensible and Scalable Framework for Decision Management in New-age Energy Systems
by: Roy, Millend, et al.
Published: (2025)
by: Roy, Millend, et al.
Published: (2025)
Panchromatic‐Guided Self‐Supervised Framework for Dual‐Camera Compressive Spectral Imaging
by: Wenzhen Lai, et al.
Published: (2026)
by: Wenzhen Lai, et al.
Published: (2026)
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Parsing of Research Documents into XML Using Formal Grammars
by: Opeoluwa Iwashokun, et al.
Published: (2024)
by: Opeoluwa Iwashokun, et al.
Published: (2024)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
by: Wang, Shouren, et al.
Published: (2025)
by: Wang, Shouren, et al.
Published: (2025)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
SELF: Self-Extend the Context Length With Logistic Growth Function
by: Dang, Phat Thanh, et al.
Published: (2025)
by: Dang, Phat Thanh, et al.
Published: (2025)
A large language model-type architecture for high-dimensional molecular potential energy surfaces
by: Zhu, Xiao, et al.
Published: (2024)
by: Zhu, Xiao, et al.
Published: (2024)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025)
by: Srinivasan, Tejas, et al.
Published: (2025)
Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution
by: Zhang, Weixing, et al.
Published: (2026)
by: Zhang, Weixing, et al.
Published: (2026)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
by: Ferrer, Robinson, et al.
Published: (2026)
by: Ferrer, Robinson, et al.
Published: (2026)
InteracTalker: Prompt-Based Human-Object Interaction with Co-Speech Gesture Generation
by: Rajan, Sreehari, et al.
Published: (2025)
by: Rajan, Sreehari, et al.
Published: (2025)
Rethinking Vision Transformer Depth via Structural Reparameterization
by: Zhou, Chengwei, et al.
Published: (2025)
by: Zhou, Chengwei, et al.
Published: (2025)
When Trust is Zero Sum: Automation Threat to Epistemic Agency
by: Malone, Emmie, et al.
Published: (2024)
by: Malone, Emmie, et al.
Published: (2024)
When Large Language Models Meet Optical Networks: Paving the Way for Automation
by: Wang, Danshi, et al.
Published: (2024)
by: Wang, Danshi, et al.
Published: (2024)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Task Calibration: Calibrating Large Language Models on Inference Tasks
by: Li, Yingjie, et al.
Published: (2024)
by: Li, Yingjie, et al.
Published: (2024)
Quantum circuit and mapping algorithms for wavepacket dynamics: case study of anharmonic hydrogen bonds in protonated and hydroxide water clusters
by: Saha, Debadrita, et al.
Published: (2024)
by: Saha, Debadrita, et al.
Published: (2024)
Similar Items
-
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning
by: Ganguly, Debargha, et al.
Published: (2024) -
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
by: Zhang, Biyao, et al.
Published: (2025) -
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
by: Ganguly, Debargha, et al.
Published: (2025) -
Trust The Typical
by: Ganguly, Debargha, et al.
Published: (2026) -
LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection
by: Zhang, Xuecen, et al.
Published: (2026)