The Attentional White Bear Effect in Transformer Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramnauth, Rebecca, Scassellati, Brian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
More than Chit-Chat: Developing Robots for Small-Talk Interactions
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024)
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
A Grounded Observer Framework for Establishing Guardrails for Foundation Models in Socially Sensitive Domains
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024)
A Robot-Assisted Approach to Small Talk Training for Adults with ASD
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2025)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2025)
Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
von: Mann, Logan, et al.
Veröffentlicht: (2025)
von: Mann, Logan, et al.
Veröffentlicht: (2025)
Gaze Behavior During a Long-Term, In-Home, Social Robot Intervention for Children with ASD
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2025)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2025)
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
Attention Sinks in Diffusion Language Models
von: Rulli, Maximo Eduardo, et al.
Veröffentlicht: (2025)
von: Rulli, Maximo Eduardo, et al.
Veröffentlicht: (2025)
Memorization in Attention-only Transformers
von: Dana, Léo, et al.
Veröffentlicht: (2024)
von: Dana, Léo, et al.
Veröffentlicht: (2024)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
Efficient Streaming Language Models with Attention Sinks
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
Cross-Attention Watermarking of Large Language Models
von: Baldassini, Folco Bertini, et al.
Veröffentlicht: (2024)
von: Baldassini, Folco Bertini, et al.
Veröffentlicht: (2024)
Attention-Aligned Reasoning for Large Language Models
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Weighted Grouped Query Attention in Transformers
von: Chinnakonduru, Sai Sena, et al.
Veröffentlicht: (2024)
von: Chinnakonduru, Sai Sena, et al.
Veröffentlicht: (2024)
Latent Multi-Head Attention for Small Language Models
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Models
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models
von: Sun, Qian, et al.
Veröffentlicht: (2024)
von: Sun, Qian, et al.
Veröffentlicht: (2024)
Word Meanings in Transformer Language Models
von: Grindrod, Jumbly, et al.
Veröffentlicht: (2025)
von: Grindrod, Jumbly, et al.
Veröffentlicht: (2025)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Efficient Attention Mechanisms for Large Language Models: A Survey
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
von: Zhao, Yida, et al.
Veröffentlicht: (2024)
von: Zhao, Yida, et al.
Veröffentlicht: (2024)
Intra-Layer Recurrence in Transformers for Language Modeling
von: Nguyen, Anthony, et al.
Veröffentlicht: (2025)
von: Nguyen, Anthony, et al.
Veröffentlicht: (2025)
Crisp Attention: Regularizing Transformers via Structured Sparsity
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
Dynamic Topic Evolution with Temporal Decay and Attention in Large Language Models
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Attention Basin: Why Contextual Position Matters in Large Language Models
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
Self-Selected Attention Span for Accelerating Large Language Model Inference
von: Jin, Tian, et al.
Veröffentlicht: (2024)
von: Jin, Tian, et al.
Veröffentlicht: (2024)
Transformer-based Causal Language Models Perform Clustering
von: Wu, Xinbo, et al.
Veröffentlicht: (2024)
von: Wu, Xinbo, et al.
Veröffentlicht: (2024)
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2024)
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2024)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
Cognitive Effects in Large Language Models
von: Shaki, Jonathan, et al.
Veröffentlicht: (2023)
von: Shaki, Jonathan, et al.
Veröffentlicht: (2023)
Transformer-based Language Models for Reasoning in the Description Logic ALCQ
von: Poulis, Angelos, et al.
Veröffentlicht: (2024)
von: Poulis, Angelos, et al.
Veröffentlicht: (2024)
Evidence of Phase Transitions in Small Transformer-Based Language Models
von: Hong, Noah, et al.
Veröffentlicht: (2025)
von: Hong, Noah, et al.
Veröffentlicht: (2025)
A Systematic Study of Compositional Syntactic Transformer Language Models
von: Zhao, Yida, et al.
Veröffentlicht: (2025)
von: Zhao, Yida, et al.
Veröffentlicht: (2025)
How Large Language Models are Transforming Machine-Paraphrased Plagiarism
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
More than Chit-Chat: Developing Robots for Small-Talk Interactions
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024) -
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026) -
A Grounded Observer Framework for Establishing Guardrails for Foundation Models in Socially Sensitive Domains
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2024) -
A Robot-Assisted Approach to Small Talk Training for Adults with ASD
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2025) -
Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
von: Mann, Logan, et al.
Veröffentlicht: (2025)