Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Jiaxin, Guan, Jian, Wang, Hongning, Wu, Wei, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback
von: Guan, Jian, et al.
Veröffentlicht: (2024)
von: Guan, Jian, et al.
Veröffentlicht: (2024)
Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models
von: Hu, Hanxu, et al.
Veröffentlicht: (2023)
von: Hu, Hanxu, et al.
Veröffentlicht: (2023)
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
Think Socially via Cognitive Reasoning
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
Data Selection via Optimal Control for Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
Codifying Natural Langauge Tasks
von: Chen, Haoyang, et al.
Veröffentlicht: (2025)
von: Chen, Haoyang, et al.
Veröffentlicht: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2025)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2025)
Language Models Hallucinate, but May Excel at Fact Verification
von: Guan, Jian, et al.
Veröffentlicht: (2023)
von: Guan, Jian, et al.
Veröffentlicht: (2023)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Towards Efficient Exact Optimization of Language Model Alignment
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
von: Ke, Pei, et al.
Veröffentlicht: (2023)
von: Ke, Pei, et al.
Veröffentlicht: (2023)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
MiniLLM: On-Policy Distillation of Large Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2023)
von: Gu, Yuxian, et al.
Veröffentlicht: (2023)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
von: Lu, Junru, et al.
Veröffentlicht: (2025)
von: Lu, Junru, et al.
Veröffentlicht: (2025)
Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
CharacterBench: Benchmarking Character Customization of Large Language Models
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
Unlocking the Potential: Benchmarking Large Language Models in Water Engineering and Research
von: Xu, Boyan, et al.
Veröffentlicht: (2024)
von: Xu, Boyan, et al.
Veröffentlicht: (2024)
SocialSim: Towards Socialized Simulation of Emotional Support Conversation
von: Chen, Zhuang, et al.
Veröffentlicht: (2025)
von: Chen, Zhuang, et al.
Veröffentlicht: (2025)
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback
von: Guan, Jian, et al.
Veröffentlicht: (2024) -
Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models
von: Hu, Hanxu, et al.
Veröffentlicht: (2023) -
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024) -
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023) -
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)