Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jin Peng, Staats, Charles, Li, Wenda, Szegedy, Christian, Weinberger, Kilian Q., Wu, Yuhuai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025)
by: Rameshkumar, Revanth, et al.
Published: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025)
by: Hassid, Michael, et al.
Published: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
Correction with Backtracking Reduces Hallucination in Summarization
by: Liu, Zhenzhen, et al.
Published: (2023)
by: Liu, Zhenzhen, et al.
Published: (2023)
On Speeding Up Language Model Evaluation
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
KELPS: A Framework for Verified Multi-Language Autoformalization via Semantic-Syntactic Alignment
by: Zhang, Jiyao, et al.
Published: (2025)
by: Zhang, Jiyao, et al.
Published: (2025)
Orchestrating LLMs with Different Personalizations
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations
by: Meng, Qian, et al.
Published: (2025)
by: Meng, Qian, et al.
Published: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
by: Qin, Yuehan, et al.
Published: (2025)
by: Qin, Yuehan, et al.
Published: (2025)
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
by: Lachenmaier, Clara, et al.
Published: (2025)
by: Lachenmaier, Clara, et al.
Published: (2025)
Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning
by: He, Yongquan, et al.
Published: (2024)
by: He, Yongquan, et al.
Published: (2024)
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
by: Jedidi, Nour, et al.
Published: (2025)
by: Jedidi, Nour, et al.
Published: (2025)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
by: Wu, Yutong, et al.
Published: (2025)
by: Wu, Yutong, et al.
Published: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
by: Liu, Jinda, et al.
Published: (2025)
by: Liu, Jinda, et al.
Published: (2025)
Language Models Don't Learn the Physical Manifestation of Language
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
A New Approach Towards Autoformalization
by: Patel, Nilay, et al.
Published: (2023)
by: Patel, Nilay, et al.
Published: (2023)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
sDPO: Don't Use Your Data All at Once
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
Faithful Autoformalization via Roundtrip Verification and Repair
by: Amrollahi, Daneshvar, et al.
Published: (2026)
by: Amrollahi, Daneshvar, et al.
Published: (2026)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
by: Bavaresco, A., et al.
Published: (2024)
by: Bavaresco, A., et al.
Published: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
by: Singh, Vikash, et al.
Published: (2026)
by: Singh, Vikash, et al.
Published: (2026)
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
by: Gong, Albert, et al.
Published: (2025)
by: Gong, Albert, et al.
Published: (2025)
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Don't Change My View: Ideological Bias Auditing in Large Language Models
by: Kröger, Paul, et al.
Published: (2025)
by: Kröger, Paul, et al.
Published: (2025)
Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents
by: Sethi, Khushal
Published: (2026)
by: Sethi, Khushal
Published: (2026)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
by: Chen, Xinxi, et al.
Published: (2024)
by: Chen, Xinxi, et al.
Published: (2024)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Learning from Synthetic Data Improves Multi-hop Reasoning
by: Kabra, Anmol, et al.
Published: (2026)
by: Kabra, Anmol, et al.
Published: (2026)
AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning
by: Bruno, Alessio
Published: (2026)
by: Bruno, Alessio
Published: (2026)
Similar Items
-
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025) -
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025) -
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025) -
Correction with Backtracking Reduces Hallucination in Summarization
by: Liu, Zhenzhen, et al.
Published: (2023) -
On Speeding Up Language Model Evaluation
by: Zhou, Jin Peng, et al.
Published: (2024)