Guardado en:
| Autores principales: | Yang, Leixin, Xiang, Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2309.12689 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations
por: Bhan, Milan, et al.
Publicado: (2024)
por: Bhan, Milan, et al.
Publicado: (2024)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
por: Baali, Massa, et al.
Publicado: (2025)
por: Baali, Massa, et al.
Publicado: (2025)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
por: Pang, Zirui, et al.
Publicado: (2025)
por: Pang, Zirui, et al.
Publicado: (2025)
Extracting Rule-based Descriptions of Attention Features in Transformers
por: Friedman, Dan, et al.
Publicado: (2025)
por: Friedman, Dan, et al.
Publicado: (2025)
Attention Smoothing Is All You Need For Unlearning
por: Zade, Saleh Zare, et al.
Publicado: (2026)
por: Zade, Saleh Zare, et al.
Publicado: (2026)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
Revisiting the Role of Label Smoothing in Enhanced Text Sentiment Classification
por: Gao, Yijie, et al.
Publicado: (2023)
por: Gao, Yijie, et al.
Publicado: (2023)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
por: Xiao, Yuxin, et al.
Publicado: (2024)
por: Xiao, Yuxin, et al.
Publicado: (2024)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
por: Jelenić, Fran, et al.
Publicado: (2023)
por: Jelenić, Fran, et al.
Publicado: (2023)
Gated Linear Attention Transformers with Hardware-Efficient Training
por: Yang, Songlin, et al.
Publicado: (2023)
por: Yang, Songlin, et al.
Publicado: (2023)
LASER: Attention with Exponential Transformation
por: Duvvuri, Sai Surya, et al.
Publicado: (2024)
por: Duvvuri, Sai Surya, et al.
Publicado: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
por: Heo, DongNyeong, et al.
Publicado: (2024)
por: Heo, DongNyeong, et al.
Publicado: (2024)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
por: Aggarwal, Shubham, et al.
Publicado: (2026)
por: Aggarwal, Shubham, et al.
Publicado: (2026)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
por: Yang, Songlin, et al.
Publicado: (2025)
por: Yang, Songlin, et al.
Publicado: (2025)
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
por: Collins, Jeremy A., et al.
Publicado: (2025)
por: Collins, Jeremy A., et al.
Publicado: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
por: Bozic, Vukasin, et al.
Publicado: (2023)
por: Bozic, Vukasin, et al.
Publicado: (2023)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
por: Ram, Dhananjay, et al.
Publicado: (2025)
por: Ram, Dhananjay, et al.
Publicado: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
por: Xiao, Da, et al.
Publicado: (2024)
por: Xiao, Da, et al.
Publicado: (2024)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
por: Tian, Ye, et al.
Publicado: (2024)
por: Tian, Ye, et al.
Publicado: (2024)
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation
por: Wang, Yu, et al.
Publicado: (2025)
por: Wang, Yu, et al.
Publicado: (2025)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
por: Nguyen, Hieu, et al.
Publicado: (2025)
por: Nguyen, Hieu, et al.
Publicado: (2025)
Automatic Differential Diagnosis using Transformer-Based Multi-Label Sequence Classification
por: Sadi, Abu Adnan, et al.
Publicado: (2024)
por: Sadi, Abu Adnan, et al.
Publicado: (2024)
Selective Attention Improves Transformer
por: Leviathan, Yaniv, et al.
Publicado: (2024)
por: Leviathan, Yaniv, et al.
Publicado: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
por: Gerami, Armin, et al.
Publicado: (2025)
por: Gerami, Armin, et al.
Publicado: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
por: Nawrot, Piotr, et al.
Publicado: (2025)
por: Nawrot, Piotr, et al.
Publicado: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
por: Chelba, Ciprian, et al.
Publicado: (2020)
por: Chelba, Ciprian, et al.
Publicado: (2020)
Selective Attention: Enhancing Transformer through Principled Context Control
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
por: Musat, Tiberiu
Publicado: (2024)
por: Musat, Tiberiu
Publicado: (2024)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
por: Brandon, William, et al.
Publicado: (2024)
por: Brandon, William, et al.
Publicado: (2024)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
por: Mihaila, George
Publicado: (2026)
por: Mihaila, George
Publicado: (2026)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
por: Dong, Yihe, et al.
Publicado: (2025)
por: Dong, Yihe, et al.
Publicado: (2025)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
por: Gerber, Isaac
Publicado: (2025)
por: Gerber, Isaac
Publicado: (2025)
Detecting and Rectifying Noisy Labels: A Similarity-based Approach
por: Huu-Tien, Dang, et al.
Publicado: (2025)
por: Huu-Tien, Dang, et al.
Publicado: (2025)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
por: Zhussip, Magauiya, et al.
Publicado: (2025)
por: Zhussip, Magauiya, et al.
Publicado: (2025)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
por: Zeris, Athanasios
Publicado: (2026)
por: Zeris, Athanasios
Publicado: (2026)
Leveraging Label Semantics and Meta-Label Refinement for Multi-Label Question Classification
por: Dong, Shi, et al.
Publicado: (2024)
por: Dong, Shi, et al.
Publicado: (2024)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
por: Chen, Feiyang, et al.
Publicado: (2025)
por: Chen, Feiyang, et al.
Publicado: (2025)
Ejemplares similares
-
Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations
por: Bhan, Milan, et al.
Publicado: (2024) -
CAARMA: Class Augmentation with Adversarial Mixup Regularization
por: Baali, Massa, et al.
Publicado: (2025) -
Label Smoothing Improves Gradient Ascent in LLM Unlearning
por: Pang, Zirui, et al.
Publicado: (2025) -
Extracting Rule-based Descriptions of Attention Features in Transformers
por: Friedman, Dan, et al.
Publicado: (2025) -
Attention Smoothing Is All You Need For Unlearning
por: Zade, Saleh Zare, et al.
Publicado: (2026)