Small Language Models Need Strong Verifiers to Self-Correct Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yunxiang, Khalifa, Muhammad, Logeswaran, Lajanugen, Kim, Jaekyeom, Lee, Moontae, Lee, Honglak, Wang, Lu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
par: Khalifa, Muhammad, et autres
Publié: (2023)
par: Khalifa, Muhammad, et autres
Publié: (2023)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
par: Khalifa, Muhammad, et autres
Publié: (2026)
par: Khalifa, Muhammad, et autres
Publié: (2026)
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025)
par: Khalifa, Muhammad, et autres
Publié: (2025)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
par: Kim, Jaekyeom, et autres
Publié: (2024)
par: Kim, Jaekyeom, et autres
Publié: (2024)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
par: Shen, Siqi, et autres
Publié: (2024)
par: Shen, Siqi, et autres
Publié: (2024)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
par: Zhang, Yunxiang, et autres
Publié: (2025)
par: Zhang, Yunxiang, et autres
Publié: (2025)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
par: Shen, Siqi, et autres
Publié: (2025)
par: Shen, Siqi, et autres
Publié: (2025)
Selective LoRA for Visual Tokens and Attention Heads
par: Luo, Tiange, et autres
Publié: (2025)
par: Luo, Tiange, et autres
Publié: (2025)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
par: Zhang, Lechen, et autres
Publié: (2024)
par: Zhang, Lechen, et autres
Publié: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
par: Zheng, Mingqian, et autres
Publié: (2023)
par: Zheng, Mingqian, et autres
Publié: (2023)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
par: Logeswaran, Lajanugen, et autres
Publié: (2026)
par: Logeswaran, Lajanugen, et autres
Publié: (2026)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
par: Zhang, Lechen, et autres
Publié: (2025)
par: Zhang, Lechen, et autres
Publié: (2025)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
par: Shu, Bangzhao, et autres
Publié: (2023)
par: Shu, Bangzhao, et autres
Publié: (2023)
Visual Test-time Scaling for GUI Agent Grounding
par: Luo, Tiange, et autres
Publié: (2025)
par: Luo, Tiange, et autres
Publié: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
par: Jang, Yunseok, et autres
Publié: (2025)
par: Jang, Yunseok, et autres
Publié: (2025)
Source-Aware Training Enables Knowledge Attribution in Language Models
par: Khalifa, Muhammad, et autres
Publié: (2024)
par: Khalifa, Muhammad, et autres
Publié: (2024)
Interactive and Expressive Code-Augmented Planning with Large Language Models
par: Liu, Anthony Z., et autres
Publié: (2024)
par: Liu, Anthony Z., et autres
Publié: (2024)
Small Language Models are Equation Reasoners
par: Kim, Bumjun, et autres
Publié: (2024)
par: Kim, Bumjun, et autres
Publié: (2024)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
par: Zhang, Yunxiang, et autres
Publié: (2025)
par: Zhang, Yunxiang, et autres
Publié: (2025)
Self-Correcting Code Generation Using Small Language Models
par: Cho, Jeonghun, et autres
Publié: (2025)
par: Cho, Jeonghun, et autres
Publié: (2025)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
par: Zhang, Yunxiang, et autres
Publié: (2025)
par: Zhang, Yunxiang, et autres
Publié: (2025)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
par: Han, Janghoon, et autres
Publié: (2025)
par: Han, Janghoon, et autres
Publié: (2025)
LitCab: Lightweight Language Model Calibration over Short- and Long-form Responses
par: Liu, Xin, et autres
Publié: (2023)
par: Liu, Xin, et autres
Publié: (2023)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
par: Lee, Seongyun, et autres
Publié: (2026)
par: Lee, Seongyun, et autres
Publié: (2026)
LG AI Research & KAIST at EHRSQL 2024: Self-Training Large Language Models with Pseudo-Labeled Unanswerable Questions for a Reliable Text-to-SQL System on EHRs
par: Jo, Yongrae, et autres
Publié: (2024)
par: Jo, Yongrae, et autres
Publié: (2024)
Do LLMs Need Inherent Reasoning Before Reinforcement Learning? A Study in Korean Self-Correction
par: Kim, Hongjin, et autres
Publié: (2026)
par: Kim, Hongjin, et autres
Publié: (2026)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
par: Lee, Hojae, et autres
Publié: (2024)
par: Lee, Hojae, et autres
Publié: (2024)
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
par: Kim, Kyuyoung, et autres
Publié: (2026)
par: Kim, Kyuyoung, et autres
Publié: (2026)
Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
par: Kim, Jiyeon, et autres
Publié: (2026)
par: Kim, Jiyeon, et autres
Publié: (2026)
Learning to Self-Verify Makes Language Models Better Reasoners
par: Chen, Yuxin, et autres
Publié: (2026)
par: Chen, Yuxin, et autres
Publié: (2026)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
par: Kim, Seoyeon, et autres
Publié: (2024)
par: Kim, Seoyeon, et autres
Publié: (2024)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
par: Khalifa, Muhammad, et autres
Publié: (2024)
par: Khalifa, Muhammad, et autres
Publié: (2024)
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks
par: Kim, Hyunjae, et autres
Publié: (2024)
par: Kim, Hyunjae, et autres
Publié: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
par: Lee, Kyungjae, et autres
Publié: (2024)
par: Lee, Kyungjae, et autres
Publié: (2024)
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
par: Kim, Jaehoon, et autres
Publié: (2025)
par: Kim, Jaehoon, et autres
Publié: (2025)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
par: Kim, Seungone, et autres
Publié: (2024)
par: Kim, Seungone, et autres
Publié: (2024)
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
par: Chae, Hyungjoo, et autres
Publié: (2025)
par: Chae, Hyungjoo, et autres
Publié: (2025)
Learning to Reason via Self-Iterative Process Feedback for Small Language Models
par: Chen, Kaiyuan, et autres
Publié: (2024)
par: Chen, Kaiyuan, et autres
Publié: (2024)
On Many-Shot In-Context Learning for Long-Context Evaluation
par: Zou, Kaijian, et autres
Publié: (2024)
par: Zou, Kaijian, et autres
Publié: (2024)
Documents similaires
-
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
par: Khalifa, Muhammad, et autres
Publié: (2023) -
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
par: Khalifa, Muhammad, et autres
Publié: (2026) -
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025) -
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
par: Kim, Jaekyeom, et autres
Publié: (2024) -
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
par: Shen, Siqi, et autres
Publié: (2024)