Saved in:
| Main Authors: | Heo, DongNyeong, Choi, Heeyoul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.15578 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shared Latent Space by Both Languages in Non-Autoregressive Neural Machine Translation
by: Heo, DongNyeong, et al.
Published: (2023)
by: Heo, DongNyeong, et al.
Published: (2023)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
by: Heo, DongNyeong, et al.
Published: (2022)
by: Heo, DongNyeong, et al.
Published: (2022)
Sentence Curve Language Models
by: Heo, DongNyeong, et al.
Published: (2026)
by: Heo, DongNyeong, et al.
Published: (2026)
N-gram Prediction and Word Difference Representations for Language Modeling
by: Heo, DongNyeong, et al.
Published: (2024)
by: Heo, DongNyeong, et al.
Published: (2024)
Dynamic Preference Multi-Objective Reinforcement Learning for Internet Network Management
by: Heo, DongNyeong, et al.
Published: (2025)
by: Heo, DongNyeong, et al.
Published: (2025)
Enhanced Labeling Technique for Reddit Text and Fine-Tuned Longformer Models for Classifying Depression Severity in English and Luganda
by: Kimera, Richard, et al.
Published: (2024)
by: Kimera, Richard, et al.
Published: (2024)
Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda
by: Kimera, Richard, et al.
Published: (2025)
by: Kimera, Richard, et al.
Published: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Probabilistic Topic Modelling with Transformer Representations
by: Reuter, Arik, et al.
Published: (2024)
by: Reuter, Arik, et al.
Published: (2024)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
by: Nam, KiHyun, et al.
Published: (2025)
by: Nam, KiHyun, et al.
Published: (2025)
Unifying Linear-Time Attention via Latent Probabilistic Modelling
by: Dolga, Rares, et al.
Published: (2024)
by: Dolga, Rares, et al.
Published: (2024)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
by: Dong, Yihe, et al.
Published: (2025)
by: Dong, Yihe, et al.
Published: (2025)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
Empirical Study of Symmetrical Reasoning in Conversational Chatbots
by: Rim, Daniela N., et al.
Published: (2024)
by: Rim, Daniela N., et al.
Published: (2024)
Convolutional Lie Operator for Sentence Classification
by: Rim, Daniela N., et al.
Published: (2025)
by: Rim, Daniela N., et al.
Published: (2025)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
by: Kang, Yanming, et al.
Published: (2023)
by: Kang, Yanming, et al.
Published: (2023)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
by: Aggarwal, Shubham, et al.
Published: (2026)
by: Aggarwal, Shubham, et al.
Published: (2026)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
by: Koh, Minsu, et al.
Published: (2025)
by: Koh, Minsu, et al.
Published: (2025)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
by: Bianchessi, Arthur S., et al.
Published: (2025)
by: Bianchessi, Arthur S., et al.
Published: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Limitations of Normalization in Attention Mechanism
by: Mudarisov, Timur, et al.
Published: (2025)
by: Mudarisov, Timur, et al.
Published: (2025)
Selective Attention Improves Transformer
by: Leviathan, Yaniv, et al.
Published: (2024)
by: Leviathan, Yaniv, et al.
Published: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
by: Lee, Tzu-Yun, et al.
Published: (2025)
by: Lee, Tzu-Yun, et al.
Published: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
by: Guo, Zhenyu, et al.
Published: (2025)
by: Guo, Zhenyu, et al.
Published: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
by: Yang, Leixin, et al.
Published: (2023)
by: Yang, Leixin, et al.
Published: (2023)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
Similar Items
-
Shared Latent Space by Both Languages in Non-Autoregressive Neural Machine Translation
by: Heo, DongNyeong, et al.
Published: (2023) -
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
by: Heo, DongNyeong, et al.
Published: (2022) -
Sentence Curve Language Models
by: Heo, DongNyeong, et al.
Published: (2026) -
N-gram Prediction and Word Difference Representations for Language Modeling
by: Heo, DongNyeong, et al.
Published: (2024) -
Dynamic Preference Multi-Objective Reinforcement Learning for Internet Network Management
by: Heo, DongNyeong, et al.
Published: (2025)