Revealing the Attention Floating Mechanism in Masked Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Xin, Huang, Pengcheng, Liu, Zhenghao, Wang, Shuo, Yan, Yukun, Xiao, Chaojun, Gu, Yu, Yu, Ge, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empirical Analysis of Decoding Biases in Masked Diffusion Models
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
by: Duan, Shaohua, et al.
Published: (2025)
by: Duan, Shaohua, et al.
Published: (2025)
ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance
by: Yao, Sijia, et al.
Published: (2025)
by: Yao, Sijia, et al.
Published: (2025)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026)
by: Gao, Xingjie, et al.
Published: (2026)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
by: Liu, Zhenghao, et al.
Published: (2025)
by: Liu, Zhenghao, et al.
Published: (2025)
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
by: Yang, Weiqing, et al.
Published: (2024)
by: Yang, Weiqing, et al.
Published: (2024)
Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains
by: Xiong, Yuqi, et al.
Published: (2026)
by: Xiong, Yuqi, et al.
Published: (2026)
Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
by: Yu, Shi, et al.
Published: (2024)
by: Yu, Shi, et al.
Published: (2024)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
by: Ma, Maotian, et al.
Published: (2026)
by: Ma, Maotian, et al.
Published: (2026)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
by: Wang, Hanbin, et al.
Published: (2023)
by: Wang, Hanbin, et al.
Published: (2023)
ADMFormer: An Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention for Traffic Forecasting
by: Gu, Ruiwen, et al.
Published: (2026)
by: Gu, Ruiwen, et al.
Published: (2026)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
by: Yang, Qing, et al.
Published: (2025)
by: Yang, Qing, et al.
Published: (2025)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
by: Li, Yishan, et al.
Published: (2026)
by: Li, Yishan, et al.
Published: (2026)
Motion Guided Token Compression for Efficient Masked Video Modeling
by: Feng, Yukun, et al.
Published: (2024)
by: Feng, Yukun, et al.
Published: (2024)
LMILAtt: A Deep Learning Model for Depression Detection from Social Media Users Enhanced by Multi-Instance Learning Based on Attention Mechanism
by: Yang, Yukun
Published: (2025)
by: Yang, Yukun
Published: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
by: Gao, Cheng, et al.
Published: (2026)
by: Gao, Cheng, et al.
Published: (2026)
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
by: Xin, Haidong, et al.
Published: (2026)
by: Xin, Haidong, et al.
Published: (2026)
CHGNN: A Semi-Supervised Contrastive Hypergraph Learning Network
by: Song, Yumeng, et al.
Published: (2023)
by: Song, Yumeng, et al.
Published: (2023)
Multi-Evidence based Fact Verification via A Confidential Graph Neural Network
by: Lan, Yuqing, et al.
Published: (2024)
by: Lan, Yuqing, et al.
Published: (2024)
Modeling User Viewing Flow Using Large Language Models for Article Recommendation
by: Liu, Zhenghao, et al.
Published: (2023)
by: Liu, Zhenghao, et al.
Published: (2023)
On the Trainability of Masked Diffusion Language Models via Blockwise Locality
by: Wang, Yuxiang, et al.
Published: (2026)
by: Wang, Yuxiang, et al.
Published: (2026)
Unifying Deductive and Abductive Reasoning in Knowledge Graphs with Masked Diffusion Model
by: Gao, Yisen, et al.
Published: (2025)
by: Gao, Yisen, et al.
Published: (2025)
RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
by: Wang, Xing, et al.
Published: (2025)
by: Wang, Xing, et al.
Published: (2025)
Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling
by: Zheng, Kaiwen, et al.
Published: (2024)
by: Zheng, Kaiwen, et al.
Published: (2024)
Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow
by: Zhong, Yangyang, et al.
Published: (2026)
by: Zhong, Yangyang, et al.
Published: (2026)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
by: Zhang, Hanzhi, et al.
Published: (2025)
by: Zhang, Hanzhi, et al.
Published: (2025)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models
by: Wei, Chenxing, et al.
Published: (2026)
by: Wei, Chenxing, et al.
Published: (2026)
Similar Items
-
Empirical Analysis of Decoding Biases in Masked Diffusion Models
by: Huang, Pengcheng, et al.
Published: (2025) -
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
by: Duan, Shaohua, et al.
Published: (2025) -
ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance
by: Yao, Sijia, et al.
Published: (2025) -
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026) -
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
by: Huang, Pengcheng, et al.
Published: (2025)