Saved in:
| Main Authors: | Behjati, Melika, Henderson, James |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2102.01223 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
by: Behjati, Melika, et al.
Published: (2025)
by: Behjati, Melika, et al.
Published: (2025)
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025)
by: Honarmand, Melika, et al.
Published: (2025)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025)
by: Wong, Liang Ze
Published: (2025)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Slot Machines: How LLMs Keep Track of Multiple Entities
by: Bogdan, Paul C., et al.
Published: (2026)
by: Bogdan, Paul C., et al.
Published: (2026)
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
by: Filvantorkaman, Melika, et al.
Published: (2026)
by: Filvantorkaman, Melika, et al.
Published: (2026)
Empirical Capacity Model for Self-Attention Neural Networks
by: Härmä, Aki, et al.
Published: (2024)
by: Härmä, Aki, et al.
Published: (2024)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
by: Dong, Yihe, et al.
Published: (2025)
by: Dong, Yihe, et al.
Published: (2025)
Neural Sequence-to-Sequence Modeling with Attention by Leveraging Deep Learning Architectures for Enhanced Contextual Understanding in Abstractive Text Summarization
by: Challagundla, Bhavith Chandra, et al.
Published: (2024)
by: Challagundla, Bhavith Chandra, et al.
Published: (2024)
Native Hybrid Attention for Efficient Sequence Modeling
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
DNAZEN: Enhanced Gene Sequence Representations via Mixed Granularities of Coding Units
by: Mao, Lei, et al.
Published: (2025)
by: Mao, Lei, et al.
Published: (2025)
Enhancing Air Quality Monitoring: A Brief Review of Federated Learning Advances
by: Yarham, Sara, et al.
Published: (2025)
by: Yarham, Sara, et al.
Published: (2025)
Single Character Perturbations Break LLM Alignment
by: Lin, Leon, et al.
Published: (2024)
by: Lin, Leon, et al.
Published: (2024)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
Coupled Query-Key Dynamics for Attention
by: Gahtan, Barak, et al.
Published: (2026)
by: Gahtan, Barak, et al.
Published: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
by: Zhao, Hangyue, et al.
Published: (2026)
by: Zhao, Hangyue, et al.
Published: (2026)
Quantum Attention by Overlap Interference: Predicting Sequences from Classical and Many-Body Quantum Data
by: Pecilli, Alessio, et al.
Published: (2026)
by: Pecilli, Alessio, et al.
Published: (2026)
CharED: Character-wise Ensemble Decoding for Large Language Models
by: Gu, Kevin, et al.
Published: (2024)
by: Gu, Kevin, et al.
Published: (2024)
Masked Gated Linear Unit
by: Tajima, Yukito, et al.
Published: (2025)
by: Tajima, Yukito, et al.
Published: (2025)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
Structured Style-Rewrite with Chain-of-Thought Planning for Low-Resource Character Dialogue
by: Zhu, Chanhui
Published: (2026)
by: Zhu, Chanhui
Published: (2026)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
by: Reddy, Natesh, et al.
Published: (2025)
by: Reddy, Natesh, et al.
Published: (2025)
Attention-Based Sampler for Diffusion Language Models
by: Zhou, Yuyan, et al.
Published: (2026)
by: Zhou, Yuyan, et al.
Published: (2026)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification
by: Lee, Hui, et al.
Published: (2025)
by: Lee, Hui, et al.
Published: (2025)
On the Representational Capacity of Recurrent Neural Language Models
by: Nowak, Franz, et al.
Published: (2023)
by: Nowak, Franz, et al.
Published: (2023)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
by: Rom, Aviad, et al.
Published: (2024)
by: Rom, Aviad, et al.
Published: (2024)
Reduction of Supervision for Biomedical Knowledge Discovery
by: Theodoropoulos, Christos, et al.
Published: (2025)
by: Theodoropoulos, Christos, et al.
Published: (2025)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
TESS: Text-to-Text Self-Conditioned Simplex Diffusion
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2023)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2023)
Similar Items
-
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
by: Behjati, Melika, et al.
Published: (2025) -
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025) -
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025) -
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024) -
Slot Machines: How LLMs Keep Track of Multiple Entities
by: Bogdan, Paul C., et al.
Published: (2026)