Learning to Reason Across Parallel Samples for LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, Jianing, Ye, Xi, Tang, Hao, Zhu, Zhigang, Choi, Eunsol |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
di: Lee, Yoonsang, et al.
Pubblicazione: (2024)
di: Lee, Yoonsang, et al.
Pubblicazione: (2024)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
di: Qi, Jianing, et al.
Pubblicazione: (2024)
di: Qi, Jianing, et al.
Pubblicazione: (2024)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
Crafting In-context Examples according to LMs' Parametric Knowledge
di: Lee, Yoonsang, et al.
Pubblicazione: (2023)
di: Lee, Yoonsang, et al.
Pubblicazione: (2023)
Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads
di: Patel, Shaswat, et al.
Pubblicazione: (2026)
di: Patel, Shaswat, et al.
Pubblicazione: (2026)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
di: Wu, Tong, et al.
Pubblicazione: (2025)
di: Wu, Tong, et al.
Pubblicazione: (2025)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
No Single Best Model for Diversity: Learning a Router for Sample Diversity
di: Liu, Yuhan, et al.
Pubblicazione: (2026)
di: Liu, Yuhan, et al.
Pubblicazione: (2026)
Textless Speech-to-Speech Translation With Limited Parallel Data
di: Diwan, Anuj, et al.
Pubblicazione: (2023)
di: Diwan, Anuj, et al.
Pubblicazione: (2023)
General-Reasoner: Advancing LLM Reasoning Across All Domains
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
di: Wang, Victor, et al.
Pubblicazione: (2025)
di: Wang, Victor, et al.
Pubblicazione: (2025)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
A Survey on Parallel Reasoning
di: Wang, Ziqi, et al.
Pubblicazione: (2025)
di: Wang, Ziqi, et al.
Pubblicazione: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2024)
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2024)
Learning Representations for Reasoning: Generalizing Across Diverse Structures
di: Zhu, Zhaocheng
Pubblicazione: (2024)
di: Zhu, Zhaocheng
Pubblicazione: (2024)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
Mitigating Temporal Misalignment by Discarding Outdated Facts
di: Zhang, Michael J. Q., et al.
Pubblicazione: (2023)
di: Zhang, Michael J. Q., et al.
Pubblicazione: (2023)
RefreshKV: Updating Small KV Cache During Long-form Generation
di: Xu, Fangyuan, et al.
Pubblicazione: (2024)
di: Xu, Fangyuan, et al.
Pubblicazione: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
di: Chen, Hung-Ting, et al.
Pubblicazione: (2024)
di: Chen, Hung-Ting, et al.
Pubblicazione: (2024)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
di: Zhang, Hongzhi, et al.
Pubblicazione: (2025)
di: Zhang, Hongzhi, et al.
Pubblicazione: (2025)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
Learning Adaptive Parallel Reasoning with Language Models
di: Pan, Jiayi, et al.
Pubblicazione: (2025)
di: Pan, Jiayi, et al.
Pubblicazione: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
di: Wan, Guangya, et al.
Pubblicazione: (2024)
di: Wan, Guangya, et al.
Pubblicazione: (2024)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
di: Bi, Jiaxi, et al.
Pubblicazione: (2026)
di: Bi, Jiaxi, et al.
Pubblicazione: (2026)
OckBench: Measuring the Efficiency of LLM Reasoning
di: Du, Zheng, et al.
Pubblicazione: (2025)
di: Du, Zheng, et al.
Pubblicazione: (2025)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
di: Lake, Thom, et al.
Pubblicazione: (2024)
di: Lake, Thom, et al.
Pubblicazione: (2024)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
di: Li, Mo, et al.
Pubblicazione: (2024)
di: Li, Mo, et al.
Pubblicazione: (2024)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
di: Zhang, Kongcheng, et al.
Pubblicazione: (2025)
di: Zhang, Kongcheng, et al.
Pubblicazione: (2025)
Boosting Language Models Reasoning with Chain-of-Knowledge Prompting
di: Wang, Jianing, et al.
Pubblicazione: (2023)
di: Wang, Jianing, et al.
Pubblicazione: (2023)
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
di: Xiong, Juming, et al.
Pubblicazione: (2026)
di: Xiong, Juming, et al.
Pubblicazione: (2026)
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
Rhapsody: A Dataset for Highlight Detection in Podcasts
di: Park, Younghan, et al.
Pubblicazione: (2025)
di: Park, Younghan, et al.
Pubblicazione: (2025)
On Language Models' Sensitivity to Suspicious Coincidences
di: Padmanabhan, Sriram, et al.
Pubblicazione: (2025)
di: Padmanabhan, Sriram, et al.
Pubblicazione: (2025)
Learning from Peers in Reasoning Models
di: Luo, Tongxu, et al.
Pubblicazione: (2025)
di: Luo, Tongxu, et al.
Pubblicazione: (2025)
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit
di: Yi, Biao, et al.
Pubblicazione: (2025)
di: Yi, Biao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
di: Lee, Yoonsang, et al.
Pubblicazione: (2024) -
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
di: Qi, Jianing, et al.
Pubblicazione: (2024) -
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025) -
Crafting In-context Examples according to LMs' Parametric Knowledge
di: Lee, Yoonsang, et al.
Pubblicazione: (2023) -
Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads
di: Patel, Shaswat, et al.
Pubblicazione: (2026)