Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Qiang, Song, Wuganjing, Lin, Zhenzhou, Chen, Feifan, Cai, Qiaolong, Li, Chen, Sui, Yongduo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
CodeGraph: Enhancing Graph Reasoning of LLMs with Code
by: Cai, Qiaolong, et al.
Published: (2024)
by: Cai, Qiaolong, et al.
Published: (2024)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025)
by: Radvand, Tara, et al.
Published: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG
by: Wang, Xihang, et al.
Published: (2026)
by: Wang, Xihang, et al.
Published: (2026)
A Little Confidence Goes a Long Way
by: Scoville, John, et al.
Published: (2024)
by: Scoville, John, et al.
Published: (2024)
Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback
by: Bai, Yu, et al.
Published: (2024)
by: Bai, Yu, et al.
Published: (2024)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
by: Cheng, Yiruo, et al.
Published: (2024)
by: Cheng, Yiruo, et al.
Published: (2024)
Unleashing the Power of Large Language Model for Denoising Recommendation
by: Wang, Shuyao, et al.
Published: (2025)
by: Wang, Shuyao, et al.
Published: (2025)
Quantification and Validation for Degree of Understanding in M2M Semantic Communications
by: Xia, Linhan, et al.
Published: (2024)
by: Xia, Linhan, et al.
Published: (2024)
Explainable Recommendation with Simulated Human Feedback
by: Tang, Jiakai, et al.
Published: (2025)
by: Tang, Jiakai, et al.
Published: (2025)
Optimal Multi-bit Generative Watermarking Schemes Under Worst-Case False-Alarm Constraints
by: Huang, Yu-Shin, et al.
Published: (2026)
by: Huang, Yu-Shin, et al.
Published: (2026)
On the Reasoning Capacity of AI Models and How to Quantify It
by: Radha, Santosh Kumar, et al.
Published: (2025)
by: Radha, Santosh Kumar, et al.
Published: (2025)
FDD CSI Feedback under Finite Downlink Training: A Rate-Distortion Perspective
by: Chen, Shuao, et al.
Published: (2026)
by: Chen, Shuao, et al.
Published: (2026)
Cost-aware LLM-based Online Dataset Annotation
by: Elumar, Eray Can, et al.
Published: (2025)
by: Elumar, Eray Can, et al.
Published: (2025)
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
by: Yong, Xixian, et al.
Published: (2025)
by: Yong, Xixian, et al.
Published: (2025)
Multi-Head Finite-State Dimension
by: Huang, Xiang, et al.
Published: (2025)
by: Huang, Xiang, et al.
Published: (2025)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
by: Ma, Huidong, et al.
Published: (2026)
by: Ma, Huidong, et al.
Published: (2026)
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
by: Ahmed, Farhan, et al.
Published: (2026)
by: Ahmed, Farhan, et al.
Published: (2026)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
HILL: Hierarchy-aware Information Lossless Contrastive Learning for Hierarchical Text Classification
by: Zhu, He, et al.
Published: (2024)
by: Zhu, He, et al.
Published: (2024)
Measuring Grammatical Diversity from Small Corpora: Derivational Entropy Rates, Mean Length of Utterances, and Annotation Invariance
by: Martin, Fermin Moscoso del Prado
Published: (2024)
by: Martin, Fermin Moscoso del Prado
Published: (2024)
Language Evolution for Evading Social Media Regulation via LLM-based Multi-agent Simulation
by: Cai, Jinyu, et al.
Published: (2024)
by: Cai, Jinyu, et al.
Published: (2024)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
by: Hu, Dou, et al.
Published: (2025)
by: Hu, Dou, et al.
Published: (2025)
HopRAG: Multi-Hop Reasoning for Logic-Aware Retrieval-Augmented Generation
by: Liu, Hao, et al.
Published: (2025)
by: Liu, Hao, et al.
Published: (2025)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
by: Elias, Noel, et al.
Published: (2024)
by: Elias, Noel, et al.
Published: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Little Giants: Synthesizing High-Quality Embedding Data at Scale
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
Reasoning and Tools for Human-Level Forecasting
by: Hsieh, Elvis, et al.
Published: (2024)
by: Hsieh, Elvis, et al.
Published: (2024)
Joint Localization and Orientation with Triple-Beam Fingerprints in Massive MIMO-OFDM
by: Zhao, Yu, et al.
Published: (2026)
by: Zhao, Yu, et al.
Published: (2026)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
by: Song, Jialin, et al.
Published: (2026)
by: Song, Jialin, et al.
Published: (2026)
LLM Reasoning Engine: Specialized Training for Enhanced Mathematical Reasoning
by: Chen, Shuguang, et al.
Published: (2024)
by: Chen, Shuguang, et al.
Published: (2024)
Integrating Pre-Trained Language Model with Physical Layer Communications
by: Lee, Ju-Hyung, et al.
Published: (2024)
by: Lee, Ju-Hyung, et al.
Published: (2024)
ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification
by: Wang, Zhensheng, et al.
Published: (2026)
by: Wang, Zhensheng, et al.
Published: (2026)
How Much Can RAG Help the Reasoning of LLM?
by: Liu, Jingyu, et al.
Published: (2024)
by: Liu, Jingyu, et al.
Published: (2024)
On Polar Coding with Feedback
by: Liu, Ling, et al.
Published: (2026)
by: Liu, Ling, et al.
Published: (2026)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
On Secrecy Capacity of Binary Beampointing Channels with Block Memory and Feedback
by: Li, Siyao, et al.
Published: (2025)
by: Li, Siyao, et al.
Published: (2025)
Clarifying orthography: Orthographic transparency as compressibility
by: Torres, Charles J., et al.
Published: (2025)
by: Torres, Charles J., et al.
Published: (2025)
Similar Items
-
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
by: Shi, Yucheng, et al.
Published: (2025) -
CodeGraph: Enhancing Graph Reasoning of LLMs with Code
by: Cai, Qiaolong, et al.
Published: (2024) -
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026) -
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025) -
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)