Saved in:
| Main Authors: | Nie, Shuaiyi, Ding, Siyu, Zhang, Wenyuan, Yu, Linhao, Yang, Tianmeng, Chen, Yao, Yin, Weichong, Sun, Yu, Wu, Hua, Liu, Tingwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.09953 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
by: Yu, Linhao, et al.
Published: (2026)
by: Yu, Linhao, et al.
Published: (2026)
Exploring the System 1 Thinking Capability of Large Reasoning Models
by: Zhang, Wenyuan, et al.
Published: (2025)
by: Zhang, Wenyuan, et al.
Published: (2025)
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
ExpSeek: Self-Triggered Experience Seeking for Web Agents
by: Zhang, Wenyuan, et al.
Published: (2026)
by: Zhang, Wenyuan, et al.
Published: (2026)
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
by: Zhang, Mengyu, et al.
Published: (2025)
by: Zhang, Mengyu, et al.
Published: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision
by: Sun, Lingzhuang, et al.
Published: (2026)
by: Sun, Lingzhuang, et al.
Published: (2026)
Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking
by: Zhang, Zefeng, et al.
Published: (2024)
by: Zhang, Zefeng, et al.
Published: (2024)
DepWiGNN: A Depth-wise Graph Neural Network for Multi-hop Spatial Reasoning in Text
by: Li, Shuaiyi, et al.
Published: (2023)
by: Li, Shuaiyi, et al.
Published: (2023)
Weights-Rotated Preference Optimization for Large Language Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Unlocking the Power of Large Language Models for Multi-table Entity Matching
by: Tang, Yingkai, et al.
Published: (2026)
by: Tang, Yingkai, et al.
Published: (2026)
MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation
by: Yuan, Zheng, et al.
Published: (2026)
by: Yuan, Zheng, et al.
Published: (2026)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
by: Zhang, Wenyuan, et al.
Published: (2025)
by: Zhang, Wenyuan, et al.
Published: (2025)
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression
by: Deng, Chenlong, et al.
Published: (2024)
by: Deng, Chenlong, et al.
Published: (2024)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
Mixture of Hidden-Dimensions Transformer
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning
by: Gao, Yaxin, et al.
Published: (2025)
by: Gao, Yaxin, et al.
Published: (2025)
Adaptive Data Augmentation for Aspect Sentiment Quad Prediction
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
by: Zhang, Zhisong, et al.
Published: (2024)
by: Zhang, Zhisong, et al.
Published: (2024)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context Reasoning
by: Zhang, Hengyu, et al.
Published: (2026)
by: Zhang, Hengyu, et al.
Published: (2026)
InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
by: Li, Shuaiyi, et al.
Published: (2025)
by: Li, Shuaiyi, et al.
Published: (2025)
MO-MIX: Multi-Objective Multi-Agent Cooperative Decision-Making With Deep Reinforcement Learning
by: Hu, Tianmeng, et al.
Published: (2026)
by: Hu, Tianmeng, et al.
Published: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
by: Liu, Haolin, et al.
Published: (2026)
by: Liu, Haolin, et al.
Published: (2026)
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
by: Zhong, Wenlin, et al.
Published: (2026)
by: Zhong, Wenlin, et al.
Published: (2026)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
by: Li, Shuaiyi, et al.
Published: (2026)
by: Li, Shuaiyi, et al.
Published: (2026)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
by: He, Xuan, et al.
Published: (2024)
by: He, Xuan, et al.
Published: (2024)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
IPS: In-Prompt Process Supervision for Short Video Content Moderation
by: Liu, Mingchao, et al.
Published: (2024)
by: Liu, Mingchao, et al.
Published: (2024)
Similar Items
-
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
by: Yu, Linhao, et al.
Published: (2026) -
Exploring the System 1 Thinking Capability of Large Reasoning Models
by: Zhang, Wenyuan, et al.
Published: (2025) -
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
by: Chen, Yao, et al.
Published: (2026) -
ExpSeek: Self-Triggered Experience Seeking for Web Agents
by: Zhang, Wenyuan, et al.
Published: (2026) -
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
by: Chen, Yao, et al.
Published: (2026)