Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Xue, Huiyin, Moosavi, Nafise Sadat, Aletras, Nikolaos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
di: Kennedy, Ian W., et al.
Pubblicazione: (2026)
di: Kennedy, Ian W., et al.
Pubblicazione: (2026)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
di: Permadi, Vynska Amalia, et al.
Pubblicazione: (2026)
di: Permadi, Vynska Amalia, et al.
Pubblicazione: (2026)
Incorporating Attribution Importance for Improving Faithfulness Metrics
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
di: Pastorino, Valeria, et al.
Pubblicazione: (2024)
di: Pastorino, Valeria, et al.
Pubblicazione: (2024)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
di: Mi, Maggie, et al.
Pubblicazione: (2025)
di: Mi, Maggie, et al.
Pubblicazione: (2025)
How to Leverage Digit Embeddings to Represent Numbers?
di: Sivakumar, Jasivan Alex, et al.
Pubblicazione: (2024)
di: Sivakumar, Jasivan Alex, et al.
Pubblicazione: (2024)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
di: Pandya, Mugdha, et al.
Pubblicazione: (2024)
di: Pandya, Mugdha, et al.
Pubblicazione: (2024)
Where does output diversity collapse in post-training?
di: Karouzos, Constantinos, et al.
Pubblicazione: (2026)
di: Karouzos, Constantinos, et al.
Pubblicazione: (2026)
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
di: Karouzos, Constantinos, et al.
Pubblicazione: (2026)
di: Karouzos, Constantinos, et al.
Pubblicazione: (2026)
We Need to Talk About Classification Evaluation Metrics in NLP
di: Vickers, Peter, et al.
Pubblicazione: (2024)
di: Vickers, Peter, et al.
Pubblicazione: (2024)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
di: Liu, Yiqi, et al.
Pubblicazione: (2023)
di: Liu, Yiqi, et al.
Pubblicazione: (2023)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
di: Mi, Maggie, et al.
Pubblicazione: (2024)
di: Mi, Maggie, et al.
Pubblicazione: (2024)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2026)
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2026)
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2023)
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2023)
Boundary-targeted Membership Inference Attacks on Safety Classifiers
di: Hughes, Anthony, et al.
Pubblicazione: (2026)
di: Hughes, Anthony, et al.
Pubblicazione: (2026)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
di: Saffari, Hamidreza, et al.
Pubblicazione: (2024)
di: Saffari, Hamidreza, et al.
Pubblicazione: (2024)
Vocabulary-level Memory Efficiency for Language Model Fine-tuning
di: Williams, Miles, et al.
Pubblicazione: (2023)
di: Williams, Miles, et al.
Pubblicazione: (2023)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
di: Aghaebe, Favour Yahdii, et al.
Pubblicazione: (2025)
di: Aghaebe, Favour Yahdii, et al.
Pubblicazione: (2025)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
di: Aghaebe, Favour Yahdii, et al.
Pubblicazione: (2026)
di: Aghaebe, Favour Yahdii, et al.
Pubblicazione: (2026)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
di: Eger, Steffen, et al.
Pubblicazione: (2025)
di: Eger, Steffen, et al.
Pubblicazione: (2025)
Instruction Following by Principled Boosting Attention of Large Language Models
di: Guardieiro, Vitoria, et al.
Pubblicazione: (2025)
di: Guardieiro, Vitoria, et al.
Pubblicazione: (2025)
Design Principle Transfer in Neural Architecture Search via Large Language Models
di: Zhou, Xun, et al.
Pubblicazione: (2024)
di: Zhou, Xun, et al.
Pubblicazione: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
di: Chrysostomou, George, et al.
Pubblicazione: (2023)
di: Chrysostomou, George, et al.
Pubblicazione: (2023)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
di: James, Joseph, et al.
Pubblicazione: (2026)
di: James, Joseph, et al.
Pubblicazione: (2026)
Towards Principled Design of Mixture-of-Experts Language Models under Memory and Inference Constraints
di: Liew, Seng Pei, et al.
Pubblicazione: (2026)
di: Liew, Seng Pei, et al.
Pubblicazione: (2026)
Self-calibration for Language Model Quantization and Pruning
di: Williams, Miles, et al.
Pubblicazione: (2024)
di: Williams, Miles, et al.
Pubblicazione: (2024)
How Private are Language Models in Abstractive Summarization?
di: Hughes, Anthony, et al.
Pubblicazione: (2024)
di: Hughes, Anthony, et al.
Pubblicazione: (2024)
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
di: Yamaguchi, Atsuki, et al.
Pubblicazione: (2026)
di: Yamaguchi, Atsuki, et al.
Pubblicazione: (2026)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
di: Yamaguchi, Atsuki, et al.
Pubblicazione: (2024)
di: Yamaguchi, Atsuki, et al.
Pubblicazione: (2024)
Exploring Gender Disparities in Automatic Speech Recognition Technology
di: ElGhazaly, Hend, et al.
Pubblicazione: (2025)
di: ElGhazaly, Hend, et al.
Pubblicazione: (2025)
Selective Attention: Enhancing Transformer through Principled Context Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
di: Stacey, Joe, et al.
Pubblicazione: (2026)
di: Stacey, Joe, et al.
Pubblicazione: (2026)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
A Multi-Task Text Classification Pipeline with Natural Language Explanations: A User-Centric Evaluation in Sentiment Analysis and Offensive Language Identification in Greek Tweets
di: Mylonas, Nikolaos, et al.
Pubblicazione: (2024)
di: Mylonas, Nikolaos, et al.
Pubblicazione: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
Attention-Based Sampler for Diffusion Language Models
di: Zhou, Yuyan, et al.
Pubblicazione: (2026)
di: Zhou, Yuyan, et al.
Pubblicazione: (2026)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2026)
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
di: Kennedy, Ian W., et al.
Pubblicazione: (2026) -
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
di: Permadi, Vynska Amalia, et al.
Pubblicazione: (2026) -
Incorporating Attribution Importance for Improving Faithfulness Metrics
di: Zhao, Zhixue, et al.
Pubblicazione: (2023) -
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
di: Pastorino, Valeria, et al.
Pubblicazione: (2024) -
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
di: Mi, Maggie, et al.
Pubblicazione: (2025)