Does Self-Attention Need Separate Weights in Transformers?
Fuente:
arXiv
Saved in:
| Main Authors: | Kowsher, Md, Prottasha, Nusrat Jahan, Yu, Chun-Nam, Garibay, Ozlem Ozmen, Yousefi, Niloofar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of Large Language Models using Semantic Knowledge Tuning
by: Prottasha, Nusrat Jahan, et al.
Published: (2024)
by: Prottasha, Nusrat Jahan, et al.
Published: (2024)
Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
by: Prottasha, Nusrat Jahan, et al.
Published: (2026)
by: Prottasha, Nusrat Jahan, et al.
Published: (2026)
Predicting Through Generation: Why Generation Is Better for Prediction
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
User Profile with Large Language Models: Construction, Updating, and Benchmarking
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
Propulsion: Steering LLM with Tiny Fine-Tuning
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
by: Kowsher, Md, et al.
Published: (2026)
by: Kowsher, Md, et al.
Published: (2026)
Explainable Detection of Implicit Influential Patterns in Conversations via Data Augmentation
by: Abdidizaji, Sina, et al.
Published: (2025)
by: Abdidizaji, Sina, et al.
Published: (2025)
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
by: Basher, Mohammad Jahid Ibna, et al.
Published: (2025)
by: Basher, Mohammad Jahid Ibna, et al.
Published: (2025)
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
L-TUNING: Synchronized Label Tuning for Prompt and Prefix in LLMs
by: Kowsher, Md., et al.
Published: (2023)
by: Kowsher, Md., et al.
Published: (2023)
RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
Token Trails: Navigating Contextual Depths in Conversational AI with ChatLLM
by: Kowsher, Md., et al.
Published: (2024)
by: Kowsher, Md., et al.
Published: (2024)
Changes by Butterflies: Farsighted Forecasting with Group Reservoir Transformer
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
When AI Does Science: Evaluating the Autonomous AI Scientist KOSMOS in Radiation Biology
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
by: Lia, Nusrat Jahan, et al.
Published: (2026)
by: Lia, Nusrat Jahan, et al.
Published: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning
by: Hossain, Elias, et al.
Published: (2026)
by: Hossain, Elias, et al.
Published: (2026)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025)
by: Jahan, Israt, et al.
Published: (2025)
Weighted Grouped Query Attention in Transformers
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
by: Ngong, Ivoline C., et al.
Published: (2024)
by: Ngong, Ivoline C., et al.
Published: (2024)
Contradiction to Consensus: Dual Perspective, Multi Source Retrieval Based Claim Verification with Source Level Disagreement using LLM
by: Biswas, Md Badsha, et al.
Published: (2026)
by: Biswas, Md Badsha, et al.
Published: (2026)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023)
by: Jahan, Israt, et al.
Published: (2023)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency
by: Sultan, Md Arafat, et al.
Published: (2025)
by: Sultan, Md Arafat, et al.
Published: (2025)
Detecting Racist Text in Bengali: An Ensemble Deep Learning Framework
by: Saruar, S. S., et al.
Published: (2024)
by: Saruar, S. S., et al.
Published: (2024)
Prompting Large Language Models to Detect Dementia Family Caregivers
by: Biswas, Md Badsha, et al.
Published: (2025)
by: Biswas, Md Badsha, et al.
Published: (2025)
How Much Context Does My Attention-Based ASR System Need?
by: Flynn, Robert, et al.
Published: (2023)
by: Flynn, Robert, et al.
Published: (2023)
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
by: Fu, Yujuan, et al.
Published: (2024)
by: Fu, Yujuan, et al.
Published: (2024)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
by: Gerber, Isaac
Published: (2025)
by: Gerber, Isaac
Published: (2025)
Metaphor Is Not All Attention Needs
by: Sorokoletova, Olga, et al.
Published: (2026)
by: Sorokoletova, Olga, et al.
Published: (2026)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Enhancing Bangla Language Next Word Prediction and Sentence Completion through Extended RNN with Bi-LSTM Model On N-gram Language
by: Islam, Md Robiul, et al.
Published: (2024)
by: Islam, Md Robiul, et al.
Published: (2024)
Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
by: Chowdhury, Masnun Nuha, et al.
Published: (2026)
by: Chowdhury, Masnun Nuha, et al.
Published: (2026)
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
by: Lia, Nusrat Jahan, et al.
Published: (2025)
by: Lia, Nusrat Jahan, et al.
Published: (2025)
Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
by: Khisa, Adity, et al.
Published: (2025)
by: Khisa, Adity, et al.
Published: (2025)
Self-Supervised Learning for Image Segmentation: A Comprehensive Survey
by: Akilan, Thangarajah, et al.
Published: (2025)
by: Akilan, Thangarajah, et al.
Published: (2025)
Transforming Fashion with AI: A Comparative Study of Large Language Models in Apparel Design
by: Lamia, Nusrat Jahan, et al.
Published: (2025)
by: Lamia, Nusrat Jahan, et al.
Published: (2025)
Similar Items
-
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
by: Kowsher, Md, et al.
Published: (2024) -
Parameter-Efficient Fine-Tuning of Large Language Models using Semantic Knowledge Tuning
by: Prottasha, Nusrat Jahan, et al.
Published: (2024) -
Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
by: Prottasha, Nusrat Jahan, et al.
Published: (2026) -
Predicting Through Generation: Why Generation Is Better for Prediction
by: Kowsher, Md, et al.
Published: (2025) -
FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
by: Kowsher, Md, et al.
Published: (2025)