Stochastic Parroting in Temporal Attention -- Regulating the Diagonal Sink
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hankemeier, Victoria, Schilling, Malte |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tailored Architectures for Time Series Forecasting: Evaluating Deep Learning Models on Gaussian Process-Generated Data
von: Hankemeier, Victoria, et al.
Veröffentlicht: (2025)
von: Hankemeier, Victoria, et al.
Veröffentlicht: (2025)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
von: Chen, Yihong, et al.
Veröffentlicht: (2026)
von: Chen, Yihong, et al.
Veröffentlicht: (2026)
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
von: Shaier, Sagi, et al.
Veröffentlicht: (2023)
von: Shaier, Sagi, et al.
Veröffentlicht: (2023)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
ASAP: Attention Sink Anchored Pruning
von: Lee, Jaehyuk, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyuk, et al.
Veröffentlicht: (2026)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
von: Zuhri, Zayd M. K., et al.
Veröffentlicht: (2025)
von: Zuhri, Zayd M. K., et al.
Veröffentlicht: (2025)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration
von: Ridder, Fabian, et al.
Veröffentlicht: (2026)
von: Ridder, Fabian, et al.
Veröffentlicht: (2026)
On the Existence and Behavior of Secondary Attention Sinks
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination
von: Li, Zihao
Veröffentlicht: (2023)
von: Li, Zihao
Veröffentlicht: (2023)
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
von: Jarvis, Devon, et al.
Veröffentlicht: (2026)
von: Jarvis, Devon, et al.
Veröffentlicht: (2026)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
von: Huang, Xingyue, et al.
Veröffentlicht: (2026)
von: Huang, Xingyue, et al.
Veröffentlicht: (2026)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
von: Yu, Mo, et al.
Veröffentlicht: (2025)
von: Yu, Mo, et al.
Veröffentlicht: (2025)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
von: Queipo-de-Llano, Enrique, et al.
Veröffentlicht: (2025)
von: Queipo-de-Llano, Enrique, et al.
Veröffentlicht: (2025)
Stochastic Parrots or ICU Experts? Large Language Models in Critical Care Medicine: A Scoping Review
von: Shi, Tongyue, et al.
Veröffentlicht: (2024)
von: Shi, Tongyue, et al.
Veröffentlicht: (2024)
Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
von: Clara, Gabriel, et al.
Veröffentlicht: (2025)
von: Clara, Gabriel, et al.
Veröffentlicht: (2025)
Scalable Stochastic Gradient Riemannian Langevin Dynamics in Non-Diagonal Metrics
von: Yu, Hanlin, et al.
Veröffentlicht: (2023)
von: Yu, Hanlin, et al.
Veröffentlicht: (2023)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
Detecting Generative Parroting through Overfitting Masked Autoencoders
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
von: Li, Siquan, et al.
Veröffentlicht: (2026)
von: Li, Siquan, et al.
Veröffentlicht: (2026)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
Personal Information Parroting in Language Models
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
von: Bai, Xueying, et al.
Veröffentlicht: (2024)
von: Bai, Xueying, et al.
Veröffentlicht: (2024)
Enhanced Structured State Space Models via Grouped FIR Filtering and Attention Sink Mechanisms
von: Meng, Tian, et al.
Veröffentlicht: (2024)
von: Meng, Tian, et al.
Veröffentlicht: (2024)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
Parrot: Multilingual Visual Instruction Tuning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
A Temporal Stochastic Bias Correction using a Machine Learning Attention model
von: Nivron, Omer, et al.
Veröffentlicht: (2024)
von: Nivron, Omer, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tailored Architectures for Time Series Forecasting: Evaluating Deep Learning Models on Gaussian Process-Generated Data
von: Hankemeier, Victoria, et al.
Veröffentlicht: (2025) -
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
von: Chen, Yihong, et al.
Veröffentlicht: (2026) -
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026) -
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
von: Shaier, Sagi, et al.
Veröffentlicht: (2023) -
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)