What Layers When: Learning to Skip Compute in LLMs with Residual Gates
Fuente:
arXiv
Guardado en:
| Autores principales: | Laitenberger, Filipe, Kopiczko, Dawid, Snoek, Cees G. M., Asano, Yuki M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KV Cache Steering for Controlling Frozen LLMs
por: Belitsky, Max, et al.
Publicado: (2025)
por: Belitsky, Max, et al.
Publicado: (2025)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
VeRA: Vector-based Random Matrix Adaptation
por: Kopiczko, Dawid J., et al.
Publicado: (2023)
por: Kopiczko, Dawid J., et al.
Publicado: (2023)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
por: Kopiczko, Dawid J., et al.
Publicado: (2026)
por: Kopiczko, Dawid J., et al.
Publicado: (2026)
NeoBabel: A Multilingual Open Tower for Visual Generation
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
por: Rastegar, Sarah, et al.
Publicado: (2024)
por: Rastegar, Sarah, et al.
Publicado: (2024)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
por: Xia, Heming, et al.
Publicado: (2025)
por: Xia, Heming, et al.
Publicado: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
por: Kang, Beomseok, et al.
Publicado: (2025)
por: Kang, Beomseok, et al.
Publicado: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
por: Zhang, Lihong, et al.
Publicado: (2025)
por: Zhang, Lihong, et al.
Publicado: (2025)
When LLMs Team Up: The Emergence of Collaborative Affective Computing
por: Lai, Wenna, et al.
Publicado: (2025)
por: Lai, Wenna, et al.
Publicado: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
por: Liu, Huabin, et al.
Publicado: (2025)
por: Liu, Huabin, et al.
Publicado: (2025)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
por: He, Zhuomin, et al.
Publicado: (2025)
por: He, Zhuomin, et al.
Publicado: (2025)
GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
por: Ou, Jie, et al.
Publicado: (2025)
por: Ou, Jie, et al.
Publicado: (2025)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
por: Hartman, Max, et al.
Publicado: (2025)
por: Hartman, Max, et al.
Publicado: (2025)
DynaPrompt: Dynamic Test-Time Prompt Tuning
por: Xiao, Zehao, et al.
Publicado: (2025)
por: Xiao, Zehao, et al.
Publicado: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
por: Sheng, Leheng, et al.
Publicado: (2026)
por: Sheng, Leheng, et al.
Publicado: (2026)
Beyond Model Adaptation at Test Time: A Survey
por: Xiao, Zehao, et al.
Publicado: (2024)
por: Xiao, Zehao, et al.
Publicado: (2024)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
por: Sanyal, Sunny, et al.
Publicado: (2024)
por: Sanyal, Sunny, et al.
Publicado: (2024)
Memorization and Knowledge Injection in Gated LLMs
por: Pan, Xu, et al.
Publicado: (2025)
por: Pan, Xu, et al.
Publicado: (2025)
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
por: Lu, Yu-Chen, et al.
Publicado: (2025)
por: Lu, Yu-Chen, et al.
Publicado: (2025)
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA
por: Roy, Nirmal, et al.
Publicado: (2024)
por: Roy, Nirmal, et al.
Publicado: (2024)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
por: Wang, Xiaoqiang, et al.
Publicado: (2026)
por: Wang, Xiaoqiang, et al.
Publicado: (2026)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
por: Khairi, Ammar, et al.
Publicado: (2025)
por: Khairi, Ammar, et al.
Publicado: (2025)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
por: Chen, Chishui, et al.
Publicado: (2026)
por: Chen, Chishui, et al.
Publicado: (2026)
Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
por: Zhu, Yunqi, et al.
Publicado: (2024)
por: Zhu, Yunqi, et al.
Publicado: (2024)
Better Language Models Exhibit Higher Visual Alignment
por: Ruthardt, Jona, et al.
Publicado: (2024)
por: Ruthardt, Jona, et al.
Publicado: (2024)
Beyond Semantics: Measuring Fine-Grained Emotion Preservation in Small Language Model-Based Machine Translation
por: Wisniewski, Dawid, et al.
Publicado: (2026)
por: Wisniewski, Dawid, et al.
Publicado: (2026)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
por: Zhou, Xinyu, et al.
Publicado: (2026)
por: Zhou, Xinyu, et al.
Publicado: (2026)
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
por: Zhu, Shengqi, et al.
Publicado: (2024)
por: Zhu, Shengqi, et al.
Publicado: (2024)
Hallucination Detection with the Internal Layers of LLMs
por: Preiß, Martin
Publicado: (2025)
por: Preiß, Martin
Publicado: (2025)
Spanish TrOCR: Leveraging Transfer Learning for Language Adaptation
por: Lauar, Filipe, et al.
Publicado: (2024)
por: Lauar, Filipe, et al.
Publicado: (2024)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
por: Ayoobi, Navid, et al.
Publicado: (2026)
por: Ayoobi, Navid, et al.
Publicado: (2026)
Generative Adversarial Reviews: When LLMs Become the Critic
por: Bougie, Nicolas, et al.
Publicado: (2024)
por: Bougie, Nicolas, et al.
Publicado: (2024)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
por: Chai, Jianzhe, et al.
Publicado: (2026)
por: Chai, Jianzhe, et al.
Publicado: (2026)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
por: Fang, Min, et al.
Publicado: (2025)
por: Fang, Min, et al.
Publicado: (2025)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
por: Asano, Yuya, et al.
Publicado: (2025)
por: Asano, Yuya, et al.
Publicado: (2025)
Adaptive Layer-skipping in Pre-trained LLMs
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
por: Moumen, Adel, et al.
Publicado: (2026)
por: Moumen, Adel, et al.
Publicado: (2026)
The Dark Patterns of Personalized Persuasion in Large Language Models: Exposing Persuasive Linguistic Features for Big Five Personality Traits in LLMs Responses
por: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Publicado: (2024)
por: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Publicado: (2024)
Ejemplares similares
-
KV Cache Steering for Controlling Frozen LLMs
por: Belitsky, Max, et al.
Publicado: (2025) -
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024) -
VeRA: Vector-based Random Matrix Adaptation
por: Kopiczko, Dawid J., et al.
Publicado: (2023) -
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
por: Kopiczko, Dawid J., et al.
Publicado: (2026) -
NeoBabel: A Multilingual Open Tower for Visual Generation
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)