Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Zhangyue, Sun, Qiushi, Guo, Qipeng, Zeng, Zhiyuan, Li, Xiaonan, Sun, Tianxiang, Chang, Cheng, Cheng, Qinyuan, Wang, Ding, Mou, Xiaofeng, Qiu, Xipeng, Huang, XuanJing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic and Generalizable Process Reward Modeling
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
Unified Active Retrieval for Retrieval Augmented Generation
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
von: Sun, Qiushi, et al.
Veröffentlicht: (2023)
von: Sun, Qiushi, et al.
Veröffentlicht: (2023)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
von: Li, Yuan, et al.
Veröffentlicht: (2025)
von: Li, Yuan, et al.
Veröffentlicht: (2025)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Scaling Laws for Fact Memorization of Large Language Models
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
How to Mitigate Overfitting in Weak-to-strong Generalization?
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
In-Memory Learning: A Declarative Learning Framework for Large Language Models
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
Zero-RAG: Towards Retrieval-Augmented Generation with Zero Redundant Knowledge
von: Luo, Qi, et al.
Veröffentlicht: (2025)
von: Luo, Qi, et al.
Veröffentlicht: (2025)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
von: Luo, Qi, et al.
Veröffentlicht: (2025)
von: Luo, Qi, et al.
Veröffentlicht: (2025)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)
von: Lv, Kai, et al.
Veröffentlicht: (2025)
Case2Code: Scalable Synthetic Data for Code Generation
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
von: Zhu, Qin, et al.
Veröffentlicht: (2025)
von: Zhu, Qin, et al.
Veröffentlicht: (2025)
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
Can Language Models Learn to Skip Steps?
von: Liu, Tengxiao, et al.
Veröffentlicht: (2024)
von: Liu, Tengxiao, et al.
Veröffentlicht: (2024)
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
von: Lv, Kai, et al.
Veröffentlicht: (2023)
von: Lv, Kai, et al.
Veröffentlicht: (2023)
How to Set the Learning Rate for Large-Scale Pre-training?
von: Zhou, Yunhua, et al.
Veröffentlicht: (2026)
von: Zhou, Yunhua, et al.
Veröffentlicht: (2026)
Understanding the CSR‐luxury paradox: The duality of luxury and responsibility in consumer perceptions
von: Yanqi Sun, et al.
Veröffentlicht: (2024)
von: Yanqi Sun, et al.
Veröffentlicht: (2024)
AI Can Learn Scientific Taste
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dynamic and Generalizable Process Reward Modeling
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025) -
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025) -
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025) -
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024) -
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)