Crisp Attention: Regularizing Transformers via Structured Sparsity
Fuente:
arXiv
Salvato in:
| Autori principali: | Gandhi, Sagar, Gandhi, Vishal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt Sentiment: The Catalyst for LLM Change
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
Personalized Image Generation from an Author Writing Style
di: Gandhi, Sagar, et al.
Pubblicazione: (2025)
di: Gandhi, Sagar, et al.
Pubblicazione: (2025)
MathDivide: Improved mathematical reasoning by large language models
di: Srivastava, Saksham Sahai, et al.
Pubblicazione: (2024)
di: Srivastava, Saksham Sahai, et al.
Pubblicazione: (2024)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
di: Tang, Zecheng, et al.
Pubblicazione: (2026)
di: Tang, Zecheng, et al.
Pubblicazione: (2026)
Training Versatile Coding Agents in Synthetic Environments
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
di: He-Yueya, Joy, et al.
Pubblicazione: (2024)
di: He-Yueya, Joy, et al.
Pubblicazione: (2024)
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
di: Gandhi, Kahaan, et al.
Pubblicazione: (2025)
di: Gandhi, Kahaan, et al.
Pubblicazione: (2025)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
Hallucination Stations: On Some Basic Limitations of Transformer-Based Language Models
di: Sikka, Varin, et al.
Pubblicazione: (2025)
di: Sikka, Varin, et al.
Pubblicazione: (2025)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
di: Dhakal, Prakash, et al.
Pubblicazione: (2024)
di: Dhakal, Prakash, et al.
Pubblicazione: (2024)
Regulating Branch Parallelism in LLM Serving
di: Gandhi, Swapnil, et al.
Pubblicazione: (2026)
di: Gandhi, Swapnil, et al.
Pubblicazione: (2026)
Scaling up the think-aloud method
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
Memorization in Attention-only Transformers
di: Dana, Léo, et al.
Pubblicazione: (2024)
di: Dana, Léo, et al.
Pubblicazione: (2024)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
Weighted Grouped Query Attention in Transformers
di: Chinnakonduru, Sai Sena, et al.
Pubblicazione: (2024)
di: Chinnakonduru, Sai Sena, et al.
Pubblicazione: (2024)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
Recursive Agent Optimization
di: Gandhi, Apurva, et al.
Pubblicazione: (2026)
di: Gandhi, Apurva, et al.
Pubblicazione: (2026)
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering
di: Patel, Udita, et al.
Pubblicazione: (2025)
di: Patel, Udita, et al.
Pubblicazione: (2025)
Fine-grainedly Synthesize Streaming Data Based On Large Language Models With Graph Structure Understanding For Data Sparsity
di: Zhang, Xin, et al.
Pubblicazione: (2024)
di: Zhang, Xin, et al.
Pubblicazione: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
di: Zhang, Zhengyan, et al.
Pubblicazione: (2024)
di: Zhang, Zhengyan, et al.
Pubblicazione: (2024)
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
di: Cheng, Xin, et al.
Pubblicazione: (2026)
di: Cheng, Xin, et al.
Pubblicazione: (2026)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
Training-Free Activation Sparsity in Large Language Models
di: Liu, James, et al.
Pubblicazione: (2024)
di: Liu, James, et al.
Pubblicazione: (2024)
The Attentional White Bear Effect in Transformer Language Models
di: Ramnauth, Rebecca, et al.
Pubblicazione: (2026)
di: Ramnauth, Rebecca, et al.
Pubblicazione: (2026)
Stream of Search (SoS): Learning to Search in Language
di: Gandhi, Kanishk, et al.
Pubblicazione: (2024)
di: Gandhi, Kanishk, et al.
Pubblicazione: (2024)
Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
End-to-End Argument Mining through Autoregressive Argumentative Structure Prediction
di: Das, Nilmadhab, et al.
Pubblicazione: (2025)
di: Das, Nilmadhab, et al.
Pubblicazione: (2025)
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
di: Bianco, Pedro Alejandro Dal, et al.
Pubblicazione: (2024)
di: Bianco, Pedro Alejandro Dal, et al.
Pubblicazione: (2024)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
di: Maliha, Maisha, et al.
Pubblicazione: (2024)
di: Maliha, Maisha, et al.
Pubblicazione: (2024)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
di: Rios, Jesus, et al.
Pubblicazione: (2025)
di: Rios, Jesus, et al.
Pubblicazione: (2025)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
di: Agarwalla, Abhinav, et al.
Pubblicazione: (2024)
di: Agarwalla, Abhinav, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Prompt Sentiment: The Catalyst for LLM Change
di: Gandhi, Vishal, et al.
Pubblicazione: (2025) -
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
di: Gandhi, Vishal, et al.
Pubblicazione: (2025) -
Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
di: Gandhi, Vishal, et al.
Pubblicazione: (2025) -
Personalized Image Generation from an Author Writing Style
di: Gandhi, Sagar, et al.
Pubblicazione: (2025) -
MathDivide: Improved mathematical reasoning by large language models
di: Srivastava, Saksham Sahai, et al.
Pubblicazione: (2024)