SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Suvarna, Ashima, Phan, Kendrick, Beikzadeh, Mehrab, Bansal, Hritik, Gabriel, Saadia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
ModelCitizens: Representing Community Voices in Online Safety
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
Adaptive Elicitation of Latent Information Using Natural Language
von: Wang, Jimmy, et al.
Veröffentlicht: (2025)
von: Wang, Jimmy, et al.
Veröffentlicht: (2025)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
Reasoning Elicitation in Language Models via Counterfactual Feedback
von: Hüyük, Alihan, et al.
Veröffentlicht: (2024)
von: Hüyük, Alihan, et al.
Veröffentlicht: (2024)
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Self-Training Elicits Concise Reasoning in Large Language Models
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
Synergy-of-Thoughts: Eliciting Efficient Reasoning in Hybrid Language Models
von: Shang, Yu, et al.
Veröffentlicht: (2024)
von: Shang, Yu, et al.
Veröffentlicht: (2024)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
Natural Language Reinforcement Learning
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
Natural Language Reinforcement Learning
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
$T^2$ of Thoughts: Temperature Tree Elicits Reasoning in Large Language Models
von: Cai, Chengkun, et al.
Veröffentlicht: (2024)
von: Cai, Chengkun, et al.
Veröffentlicht: (2024)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
Reverse Thinking Makes LLMs Stronger Reasoners
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks
von: Hegazy, Mahmood
Veröffentlicht: (2024)
von: Hegazy, Mahmood
Veröffentlicht: (2024)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
von: Chen, Keru, et al.
Veröffentlicht: (2026)
von: Chen, Keru, et al.
Veröffentlicht: (2026)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
von: Peng, Miao, et al.
Veröffentlicht: (2025)
von: Peng, Miao, et al.
Veröffentlicht: (2025)
Reinforcement Learning Enhanced LLMs: A Survey
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
von: Lu, Pan, et al.
Veröffentlicht: (2023)
von: Lu, Pan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026) -
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024) -
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
von: Bansal, Hritik, et al.
Veröffentlicht: (2025) -
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023) -
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)