Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Gengyang, Wu, Zheng-Fan, Bao, Siqi, Wu, Yunfang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy
von: Li, Gengyang, et al.
Veröffentlicht: (2025)
von: Li, Gengyang, et al.
Veröffentlicht: (2025)
SyncThink: A Training-Free Strategy to Align Inference Termination with Reasoning Saturation
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
Lost in the Passage: Passage-level In-context Learning Does Not Necessarily Need a "Passage"
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
Shared Imagination: LLMs Hallucinate Alike
von: Zhou, Yilun, et al.
Veröffentlicht: (2024)
von: Zhou, Yilun, et al.
Veröffentlicht: (2024)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
Token-Level Entropy Reveals Demographic Disparities in Language Models
von: Lee, Messi H. J.
Veröffentlicht: (2025)
von: Lee, Messi H. J.
Veröffentlicht: (2025)
What Makes Two Language Models Think Alike?
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
Ungrammatical-syntax-based In-context Example Selection for Grammatical Error Correction
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
Unsupervised Distractor Generation via Large Language Model Distilling and Counterfactual Contrastive Decoding
von: Qu, Fanyi, et al.
Veröffentlicht: (2024)
von: Qu, Fanyi, et al.
Veröffentlicht: (2024)
Composable Cross-prompt Essay Scoring by Merging Models
von: Lee, Sanwoo, et al.
Veröffentlicht: (2025)
von: Lee, Sanwoo, et al.
Veröffentlicht: (2025)
SCOI: Syntax-augmented Coverage-based In-context Example Selection for Machine Translation
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
Going Beyond Word Matching: Syntax Improves In-context Example Selection for Machine Translation
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
Evaluating the Capability of Large-scale Language Models on Chinese Grammatical Error Correction Task
von: Qu, Fanyi, et al.
Veröffentlicht: (2023)
von: Qu, Fanyi, et al.
Veröffentlicht: (2023)
SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation
von: Yang, Yutong, et al.
Veröffentlicht: (2026)
von: Yang, Yutong, et al.
Veröffentlicht: (2026)
NYK-MS: A Well-annotated Multi-modal Metaphor and Sarcasm Understanding Benchmark on Cartoon-Caption Dataset
von: Chang, Ke, et al.
Veröffentlicht: (2024)
von: Chang, Ke, et al.
Veröffentlicht: (2024)
Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive Programming
von: Zhang, Qianfan, et al.
Veröffentlicht: (2026)
von: Zhang, Qianfan, et al.
Veröffentlicht: (2026)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
von: Leong, Chak Tou, et al.
Veröffentlicht: (2024)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2024)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
Oops, Wait: Token-Level Signals as a Lens into LLM Reasoning
von: Hwang, Jaehui, et al.
Veröffentlicht: (2026)
von: Hwang, Jaehui, et al.
Veröffentlicht: (2026)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptions
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
von: Tang, Chenming, et al.
Veröffentlicht: (2024)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
ProxyAttn: Guided Sparse Attention via Representative Heads
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
von: Li, Yiyuan, et al.
Veröffentlicht: (2026)
von: Li, Yiyuan, et al.
Veröffentlicht: (2026)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
von: Liu, Runze, et al.
Veröffentlicht: (2025)
von: Liu, Runze, et al.
Veröffentlicht: (2025)
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
von: Tang, Chenming, et al.
Veröffentlicht: (2026)
von: Tang, Chenming, et al.
Veröffentlicht: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
von: Hao, Zhuoyuan, et al.
Veröffentlicht: (2026)
von: Hao, Zhuoyuan, et al.
Veröffentlicht: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy
von: Li, Gengyang, et al.
Veröffentlicht: (2025) -
SyncThink: A Training-Free Strategy to Align Inference Termination with Reasoning Saturation
von: Li, Gengyang, et al.
Veröffentlicht: (2026) -
Lost in the Passage: Passage-level In-context Learning Does Not Necessarily Need a "Passage"
von: Sun, Hao, et al.
Veröffentlicht: (2025) -
Shared Imagination: LLMs Hallucinate Alike
von: Zhou, Yilun, et al.
Veröffentlicht: (2024) -
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025)