When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Singhi, Nishad, Bansal, Hritik, Hosseini, Arian, Grover, Aditya, Chang, Kai-Wei, Rohrbach, Marcus, Rohrbach, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
by: Singhi, Nishad, et al.
Published: (2026)
by: Singhi, Nishad, et al.
Published: (2026)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
by: Bansal, Hritik, et al.
Published: (2023)
by: Bansal, Hritik, et al.
Published: (2023)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026)
by: Grebe, Jonas Henry, et al.
Published: (2026)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
Generative Verifiers: Reward Modeling as Next-Token Prediction
by: Zhang, Lunjun, et al.
Published: (2024)
by: Zhang, Lunjun, et al.
Published: (2024)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
by: AbdElhameed, Marwan, et al.
Published: (2024)
by: AbdElhameed, Marwan, et al.
Published: (2024)
Chrono: A Simple Blueprint for Representing Time in MLLMs
by: Rodriguez, Hector, et al.
Published: (2024)
by: Rodriguez, Hector, et al.
Published: (2024)
Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving
by: Shidara, Kie, et al.
Published: (2026)
by: Shidara, Kie, et al.
Published: (2026)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
by: Braun, Tobias, et al.
Published: (2026)
by: Braun, Tobias, et al.
Published: (2026)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Predicting Implicit Arguments in Procedural Video Instructions
by: Batra, Anil, et al.
Published: (2025)
by: Batra, Anil, et al.
Published: (2025)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026)
by: Rodriguez, Hector G., et al.
Published: (2026)
En búsqueda de un cuidado universal y cultural
by: Cecilia Rohrbach
Published: (2007)
by: Cecilia Rohrbach
Published: (2007)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
by: Zhou, Ruochen, et al.
Published: (2025)
by: Zhou, Ruochen, et al.
Published: (2025)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
by: Arora, Aditya, et al.
Published: (2026)
by: Arora, Aditya, et al.
Published: (2026)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)
by: Wei, Jiaqi, et al.
Published: (2026)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
by: Gao, Chang, et al.
Published: (2023)
by: Gao, Chang, et al.
Published: (2023)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025)
by: Zohrabi, Reihaneh, et al.
Published: (2025)
Design-Based Research zur Weiterentwicklung der chemiedidaktischen Lehrerausbildung zu Schülervorstellungen
by: Rohrbach-Lochner, Friederike
Published: (2022)
by: Rohrbach-Lochner, Friederike
Published: (2022)
Historic perspectives from anthropology. Reflections proposed to Transcultural Nursing
by: Cecilia Rohrbach Viadas
Published: (2015)
by: Cecilia Rohrbach Viadas
Published: (2015)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
by: Suvarna, Ashima, et al.
Published: (2026)
by: Suvarna, Ashima, et al.
Published: (2026)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers
by: Lu, Jack, et al.
Published: (2025)
by: Lu, Jack, et al.
Published: (2025)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
by: Sharan, Aditya, et al.
Published: (2026)
by: Sharan, Aditya, et al.
Published: (2026)
EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection
by: Marik, Aritra, et al.
Published: (2026)
by: Marik, Aritra, et al.
Published: (2026)
Self-consistent Reasoning For Solving Math Word Problems
by: Xiong, Jing, et al.
Published: (2022)
by: Xiong, Jing, et al.
Published: (2022)
Atiyah-Segal completion for the Hermitian K-theory of Symplectic Groups
by: Hornbostel, Jens, et al.
Published: (2023)
by: Hornbostel, Jens, et al.
Published: (2023)
Similar Items
-
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
by: Singhi, Nishad, et al.
Published: (2026) -
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024) -
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024) -
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
by: Bansal, Hritik, et al.
Published: (2023) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)