Thinking LLMs: General Instruction Following with Thought Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Tianhao, Lan, Janice, Yuan, Weizhe, Jiao, Jiantao, Weston, Jason, Sukhbaatar, Sainbayar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
di: Xiong, Wei, et al.
Pubblicazione: (2025)
di: Xiong, Wei, et al.
Pubblicazione: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
di: Yu, Ping, et al.
Pubblicazione: (2025)
di: Yu, Ping, et al.
Pubblicazione: (2025)
Iterative Reasoning Preference Optimization
di: Pang, Richard Yuanzhe, et al.
Pubblicazione: (2024)
di: Pang, Richard Yuanzhe, et al.
Pubblicazione: (2024)
Contextual Position Encoding: Learning to Count What's Important
di: Golovneva, Olga, et al.
Pubblicazione: (2024)
di: Golovneva, Olga, et al.
Pubblicazione: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
di: Xu, Jing, et al.
Pubblicazione: (2023)
di: Xu, Jing, et al.
Pubblicazione: (2023)
Reverse Training to Nurse the Reversal Curse
di: Golovneva, Olga, et al.
Pubblicazione: (2024)
di: Golovneva, Olga, et al.
Pubblicazione: (2024)
Self-Rewarding Language Models
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
Following Length Constraints in Instructions
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
Self-Challenging Language Model Agents
di: Zhou, Yifei, et al.
Pubblicazione: (2025)
di: Zhou, Yifei, et al.
Pubblicazione: (2025)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
di: Yu, Ping, et al.
Pubblicazione: (2025)
di: Yu, Ping, et al.
Pubblicazione: (2025)
Self-Consistency Preference Optimization
di: Prasad, Archiki, et al.
Pubblicazione: (2024)
di: Prasad, Archiki, et al.
Pubblicazione: (2024)
System-Level Natural Language Feedback
di: Yuan, Weizhe, et al.
Pubblicazione: (2023)
di: Yuan, Weizhe, et al.
Pubblicazione: (2023)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
di: Sukhbaatar, Sainbayar, et al.
Pubblicazione: (2024)
di: Sukhbaatar, Sainbayar, et al.
Pubblicazione: (2024)
Multi-Token Attention
di: Golovneva, Olga, et al.
Pubblicazione: (2025)
di: Golovneva, Olga, et al.
Pubblicazione: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
di: Lanchantin, Jack, et al.
Pubblicazione: (2025)
di: Lanchantin, Jack, et al.
Pubblicazione: (2025)
HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation
di: Jiao, YiHan, et al.
Pubblicazione: (2025)
di: Jiao, YiHan, et al.
Pubblicazione: (2025)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
di: Cao, Jie, et al.
Pubblicazione: (2026)
di: Cao, Jie, et al.
Pubblicazione: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
Self-Improving Pretraining: using post-trained models to pretrain better models
di: Tan, Ellen Xiaoqing, et al.
Pubblicazione: (2026)
di: Tan, Ellen Xiaoqing, et al.
Pubblicazione: (2026)
The Instruction Gap: LLMs get lost in Following Instruction
di: Tripathi, Vishesh, et al.
Pubblicazione: (2025)
di: Tripathi, Vishesh, et al.
Pubblicazione: (2025)
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking
di: Ji, Shihao, et al.
Pubblicazione: (2025)
di: Ji, Shihao, et al.
Pubblicazione: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
di: Ning, Xuefei, et al.
Pubblicazione: (2023)
di: Ning, Xuefei, et al.
Pubblicazione: (2023)
EmbedLLM: Learning Compact Representations of Large Language Models
di: Zhuang, Richard, et al.
Pubblicazione: (2024)
di: Zhuang, Richard, et al.
Pubblicazione: (2024)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
di: Su, DiJia, et al.
Pubblicazione: (2024)
di: Su, DiJia, et al.
Pubblicazione: (2024)
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
di: Xiang, Violet, et al.
Pubblicazione: (2025)
di: Xiang, Violet, et al.
Pubblicazione: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
di: Liu, Bo, et al.
Pubblicazione: (2025)
di: Liu, Bo, et al.
Pubblicazione: (2025)
Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
di: Verma, Pulkit, et al.
Pubblicazione: (2025)
di: Verma, Pulkit, et al.
Pubblicazione: (2025)
Structured Thinking Matters: Improving LLMs Generalization in Causal Inference Tasks
di: Sun, Wentao, et al.
Pubblicazione: (2025)
di: Sun, Wentao, et al.
Pubblicazione: (2025)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
Generative AI Security: Challenges and Countermeasures
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
di: Zhang, Dalong, et al.
Pubblicazione: (2025)
di: Zhang, Dalong, et al.
Pubblicazione: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
di: Ruan, Kai, et al.
Pubblicazione: (2024)
di: Ruan, Kai, et al.
Pubblicazione: (2024)
Empowering Persian LLMs for Instruction Following: A Novel Dataset and Training Approach
di: Mokhtarabadi, Hojjat, et al.
Pubblicazione: (2024)
di: Mokhtarabadi, Hojjat, et al.
Pubblicazione: (2024)
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
di: Sakai, Yusuke, et al.
Pubblicazione: (2025)
di: Sakai, Yusuke, et al.
Pubblicazione: (2025)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
di: Moore, Kyle, et al.
Pubblicazione: (2024)
di: Moore, Kyle, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024) -
StepWiser: Stepwise Generative Judges for Wiser Reasoning
di: Xiong, Wei, et al.
Pubblicazione: (2025) -
R.I.P.: Better Models by Survival of the Fittest Prompts
di: Yu, Ping, et al.
Pubblicazione: (2025) -
Iterative Reasoning Preference Optimization
di: Pang, Richard Yuanzhe, et al.
Pubblicazione: (2024) -
Contextual Position Encoding: Learning to Count What's Important
di: Golovneva, Olga, et al.
Pubblicazione: (2024)