Saved in:
| Main Authors: | Wang, Yudong, Yang, Zhe, Ma, Wenhan, Sui, Zhifang, Zhao, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.08944 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
by: Ma, Wenhan, et al.
Published: (2025)
by: Ma, Wenhan, et al.
Published: (2025)
Towards Better RL Training Data Utilization via Second-Order Rollout
by: Yang, Zhe, et al.
Published: (2026)
by: Yang, Zhe, et al.
Published: (2026)
Exploring Activation Patterns of Parameters in Language Models
by: Wang, Yudong, et al.
Published: (2024)
by: Wang, Yudong, et al.
Published: (2024)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
by: Yang, Zhe, et al.
Published: (2025)
by: Yang, Zhe, et al.
Published: (2025)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
by: Yang, Zhe, et al.
Published: (2024)
by: Yang, Zhe, et al.
Published: (2024)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
by: Yang, Zhe, et al.
Published: (2023)
by: Yang, Zhe, et al.
Published: (2023)
HistLens: Mapping Idea Change across Concepts and Corpora
by: Jing, Yi, et al.
Published: (2026)
by: Jing, Yi, et al.
Published: (2026)
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
CoLT: Reasoning with Chain of Latent Tool Calls
by: Zhu, Fangwei, et al.
Published: (2026)
by: Zhu, Fangwei, et al.
Published: (2026)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
by: Zhan, Weidong, et al.
Published: (2025)
by: Zhan, Weidong, et al.
Published: (2025)
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template Decomposition
by: Zhu, Fangwei, et al.
Published: (2024)
by: Zhu, Fangwei, et al.
Published: (2024)
Chain-of-Thought Tokens are Computer Program Variables
by: Zhu, Fangwei, et al.
Published: (2025)
by: Zhu, Fangwei, et al.
Published: (2025)
SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine
by: Wang, Xiaochen, et al.
Published: (2024)
by: Wang, Xiaochen, et al.
Published: (2024)
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
by: Yang, Zhe, et al.
Published: (2024)
by: Yang, Zhe, et al.
Published: (2024)
Language Models Encode the Value of Numbers Linearly
by: Zhu, Fangwei, et al.
Published: (2024)
by: Zhu, Fangwei, et al.
Published: (2024)
Reinforcement Pre-Training
by: Dong, Qingxiu, et al.
Published: (2025)
by: Dong, Qingxiu, et al.
Published: (2025)
FSM: A Finite State Machine Based Zero-Shot Prompting Paradigm for Multi-Hop Question Answering
by: Wang, Xiaochen, et al.
Published: (2024)
by: Wang, Xiaochen, et al.
Published: (2024)
Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
by: Meng, Xiangdi, et al.
Published: (2024)
by: Meng, Xiangdi, et al.
Published: (2024)
Plug-and-Play Training Framework for Preference Optimization
by: Ma, Jingyuan, et al.
Published: (2024)
by: Ma, Jingyuan, et al.
Published: (2024)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
by: Yang, Yixin, et al.
Published: (2025)
by: Yang, Yixin, et al.
Published: (2025)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
by: Yang, Yixin, et al.
Published: (2024)
by: Yang, Yixin, et al.
Published: (2024)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
by: Hwang, Jena D., et al.
Published: (2026)
by: Hwang, Jena D., et al.
Published: (2026)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
by: Xia, Heming, et al.
Published: (2024)
by: Xia, Heming, et al.
Published: (2024)
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA
by: Yin, Jianing, et al.
Published: (2026)
by: Yin, Jianing, et al.
Published: (2026)
DragonVerseQA: Open-Domain Long-Form Context-Aware Question-Answering
by: Lahiri, Aritra Kumar, et al.
Published: (2024)
by: Lahiri, Aritra Kumar, et al.
Published: (2024)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
ArbGraph: Conflict-Aware Evidence Arbitration for Reliable Long-Form Retrieval-Augmented Generation
by: Niu, Qingying, et al.
Published: (2026)
by: Niu, Qingying, et al.
Published: (2026)
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
by: Liu, Yudong, et al.
Published: (2025)
by: Liu, Yudong, et al.
Published: (2025)
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens
by: Luo, Weiyao, et al.
Published: (2024)
by: Luo, Weiyao, et al.
Published: (2024)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
by: Bhalerao, Parth, et al.
Published: (2026)
by: Bhalerao, Parth, et al.
Published: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
RAG-BioQA: A Retrieval-Augmented Generation Framework for Long-Form Biomedical Question Answering
by: Panchumarthi, Lovely Yeswanth, et al.
Published: (2025)
by: Panchumarthi, Lovely Yeswanth, et al.
Published: (2025)
Enhancing Long Document Long Form Summarisation with Self-Planning
by: Du, Xiaotang, et al.
Published: (2025)
by: Du, Xiaotang, et al.
Published: (2025)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
Similar Items
-
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
by: Ma, Wenhan, et al.
Published: (2025) -
Towards Better RL Training Data Utilization via Second-Order Rollout
by: Yang, Zhe, et al.
Published: (2026) -
Exploring Activation Patterns of Parameters in Language Models
by: Wang, Yudong, et al.
Published: (2024) -
A Probabilistic Inference Scaling Theory for LLM Self-Correction
by: Yang, Zhe, et al.
Published: (2025) -
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
by: Yang, Zhe, et al.
Published: (2024)