Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Guo, Zhiyu, Kamigaito, Hidetaka, Wanatnabe, Taro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multilinguality of Large Language Models From a Structural Perspective
por: Sakajo, Haruki, et al.
Publicado: (2026)
por: Sakajo, Haruki, et al.
Publicado: (2026)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
por: Sakai, Yusuke, et al.
Publicado: (2024)
por: Sakai, Yusuke, et al.
Publicado: (2024)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
por: Sakai, Yusuke, et al.
Publicado: (2024)
por: Sakai, Yusuke, et al.
Publicado: (2024)
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
por: Sakai, Yusuke, et al.
Publicado: (2023)
por: Sakai, Yusuke, et al.
Publicado: (2023)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
por: Sakai, Yusuke, et al.
Publicado: (2024)
por: Sakai, Yusuke, et al.
Publicado: (2024)
Model-based Subsampling for Knowledge Graph Completion
por: Feng, Xincan, et al.
Publicado: (2023)
por: Feng, Xincan, et al.
Publicado: (2023)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
por: Guo, Zhiyu, et al.
Publicado: (2024)
por: Guo, Zhiyu, et al.
Publicado: (2024)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
por: Fang, Gongfan, et al.
Publicado: (2024)
por: Fang, Gongfan, et al.
Publicado: (2024)
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
por: Sakai, Yusuke, et al.
Publicado: (2025)
por: Sakai, Yusuke, et al.
Publicado: (2025)
Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding
por: Natsumi, Koki, et al.
Publicado: (2025)
por: Natsumi, Koki, et al.
Publicado: (2025)
Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
por: Hashimoto, Wataru, et al.
Publicado: (2025)
por: Hashimoto, Wataru, et al.
Publicado: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
por: Zheng, Haizhong, et al.
Publicado: (2024)
por: Zheng, Haizhong, et al.
Publicado: (2024)
Tonguescape: Exploring Language Models Understanding of Vowel Articulation
por: Sakajo, Haruki, et al.
Publicado: (2025)
por: Sakajo, Haruki, et al.
Publicado: (2025)
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
por: Sakajo, Haruki, et al.
Publicado: (2026)
por: Sakajo, Haruki, et al.
Publicado: (2026)
LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models
por: Su, Yupeng, et al.
Publicado: (2024)
por: Su, Yupeng, et al.
Publicado: (2024)
GLU Attention Improve Transformer
por: Wang, Zehao
Publicado: (2025)
por: Wang, Zehao
Publicado: (2025)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
por: Akhauri, Yash, et al.
Publicado: (2024)
por: Akhauri, Yash, et al.
Publicado: (2024)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
por: Sakai, Yusuke, et al.
Publicado: (2026)
por: Sakai, Yusuke, et al.
Publicado: (2026)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
por: Sakai, Yusuke, et al.
Publicado: (2026)
por: Sakai, Yusuke, et al.
Publicado: (2026)
From Formal Language Theory to Statistical Learning: Finite Observability of Subregular Languages
por: Hayashi, Katsuhiko, et al.
Publicado: (2025)
por: Hayashi, Katsuhiko, et al.
Publicado: (2025)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
por: Xu, Peng, et al.
Publicado: (2024)
por: Xu, Peng, et al.
Publicado: (2024)
Adaptive Pruning for Large Language Models with Structural Importance Awareness
por: Zheng, Haotian, et al.
Publicado: (2024)
por: Zheng, Haotian, et al.
Publicado: (2024)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
por: Lyu, Boxuan, et al.
Publicado: (2025)
por: Lyu, Boxuan, et al.
Publicado: (2025)
Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition
por: Liu, Ziyang
Publicado: (2026)
por: Liu, Ziyang
Publicado: (2026)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
por: Li, Yun, et al.
Publicado: (2023)
por: Li, Yun, et al.
Publicado: (2023)
Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
por: Sakajo, Haruki, et al.
Publicado: (2025)
por: Sakajo, Haruki, et al.
Publicado: (2025)
Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences
por: Atuhurra, Jesse, et al.
Publicado: (2024)
por: Atuhurra, Jesse, et al.
Publicado: (2024)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
por: Nakamura, Taishi, et al.
Publicado: (2025)
por: Nakamura, Taishi, et al.
Publicado: (2025)
Unified Interpretation of Smoothing Methods for Negative Sampling Loss Functions in Knowledge Graph Embedding
por: Feng, Xincan, et al.
Publicado: (2024)
por: Feng, Xincan, et al.
Publicado: (2024)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
por: Huang, Weiyu, et al.
Publicado: (2025)
por: Huang, Weiyu, et al.
Publicado: (2025)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
por: Dinuta, Eduard Stefan, et al.
Publicado: (2025)
por: Dinuta, Eduard Stefan, et al.
Publicado: (2025)
GWQ: Gradient-Aware Weight Quantization for Large Language Models
por: Shao, Yihua, et al.
Publicado: (2024)
por: Shao, Yihua, et al.
Publicado: (2024)
Large Language Models are Pattern Matchers: Editing Semi-Structured and Structured Documents with ChatGPT
por: Weber, Irene
Publicado: (2024)
por: Weber, Irene
Publicado: (2024)
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
por: Juseon-Do, et al.
Publicado: (2024)
por: Juseon-Do, et al.
Publicado: (2024)
PAT: Pruning-Aware Tuning for Large Language Models
por: Liu, Yijiang, et al.
Publicado: (2024)
por: Liu, Yijiang, et al.
Publicado: (2024)
Align to Structure: Aligning Large Language Models with Structural Information
por: Kim, Zae Myung, et al.
Publicado: (2025)
por: Kim, Zae Myung, et al.
Publicado: (2025)
Structured Agent Distillation for Large Language Model
por: Liu, Jun, et al.
Publicado: (2025)
por: Liu, Jun, et al.
Publicado: (2025)
Ejemplares similares
-
Multilinguality of Large Language Models From a Structural Perspective
por: Sakajo, Haruki, et al.
Publicado: (2026) -
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
por: Sakai, Yusuke, et al.
Publicado: (2024) -
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
por: Hashimoto, Wataru, et al.
Publicado: (2024) -
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
por: Sakai, Yusuke, et al.
Publicado: (2024) -
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
por: Sakai, Yusuke, et al.
Publicado: (2023)