Saved in:
| Main Authors: | Xia, Xiaojie, Zhang, Huigang, Zhong, Chaoliang, Sun, Jun, Oishi, Yusuke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.11667 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
by: Feng, Cheng, et al.
Published: (2026)
by: Feng, Cheng, et al.
Published: (2026)
Native Hybrid Attention for Efficient Sequence Modeling
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
by: Ren, Yuxin, et al.
Published: (2026)
by: Ren, Yuxin, et al.
Published: (2026)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
by: Xie, Yuhan, et al.
Published: (2026)
by: Xie, Yuhan, et al.
Published: (2026)
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
GSTAM: Efficient Graph Distillation with Structural Attention-Matching
by: Rasti-Meymandi, Arash, et al.
Published: (2024)
by: Rasti-Meymandi, Arash, et al.
Published: (2024)
Enhancing Molecular Property Prediction with Auxiliary Learning and Task-Specific Adaptation
by: Dey, Vishal, et al.
Published: (2024)
by: Dey, Vishal, et al.
Published: (2024)
Efficient Multi-Task Modeling through Automated Fusion of Trained Models
by: Zhou, Jingxuan, et al.
Published: (2025)
by: Zhou, Jingxuan, et al.
Published: (2025)
EffiCANet: Efficient Time Series Forecasting with Convolutional Attention
by: Zhou, Xinxing, et al.
Published: (2024)
by: Zhou, Xinxing, et al.
Published: (2024)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
by: Saini, Dhruv, et al.
Published: (2026)
by: Saini, Dhruv, et al.
Published: (2026)
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
Efficient Multi-Task Reinforcement Learning via Task-Specific Action Correction
by: Feng, Jinyuan, et al.
Published: (2024)
by: Feng, Jinyuan, et al.
Published: (2024)
Mission-driven Exploration for Accelerated Deep Reinforcement Learning with Temporal Logic Task Specifications
by: Wang, Jun, et al.
Published: (2023)
by: Wang, Jun, et al.
Published: (2023)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
CADENT: Gated Hybrid Distillation for Sample-Efficient Transfer in Reinforcement Learning
by: Alinejad, Mahyar, et al.
Published: (2026)
by: Alinejad, Mahyar, et al.
Published: (2026)
Task-Specific Generative Dataset Distillation with Difficulty-Guided Sampling
by: Li, Mingzhuo, et al.
Published: (2025)
by: Li, Mingzhuo, et al.
Published: (2025)
Aligning Knowledge Graphs Provided by Humans and Generated from Neural Networks in Specific Tasks
by: Li, Tangrui, et al.
Published: (2024)
by: Li, Tangrui, et al.
Published: (2024)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
by: Shen, Yuhao, et al.
Published: (2026)
by: Shen, Yuhao, et al.
Published: (2026)
Easy Adaptation: An Efficient Task-Specific Knowledge Injection Method for Large Models in Resource-Constrained Environments
by: Chen, Dong, et al.
Published: (2025)
by: Chen, Dong, et al.
Published: (2025)
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
by: Fan, Tao, et al.
Published: (2025)
by: Fan, Tao, et al.
Published: (2025)
FISformer: Replacing Self-Attention with a Fuzzy Inference System in Transformer Models for Time Series Forecasting
by: Haznedar, Bulent, et al.
Published: (2026)
by: Haznedar, Bulent, et al.
Published: (2026)
Online Policy Distillation with Decision-Attention
by: Yu, Xinqiang, et al.
Published: (2024)
by: Yu, Xinqiang, et al.
Published: (2024)
Resource-Efficient Iterative LLM-Based NAS with Feedback Memory
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
by: Jackermeier, Mathias, et al.
Published: (2024)
by: Jackermeier, Mathias, et al.
Published: (2024)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context
by: Augustin, Maximilian, et al.
Published: (2024)
by: Augustin, Maximilian, et al.
Published: (2024)
Speculative Coreset Selection for Task-Specific Fine-tuning
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
Data-Efficient Symbolic Regression via Foundation Model Distillation
by: Ying, Wangyang, et al.
Published: (2025)
by: Ying, Wangyang, et al.
Published: (2025)
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
by: Li, Derek, et al.
Published: (2025)
by: Li, Derek, et al.
Published: (2025)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
ENA: Efficient N-dimensional Attention
by: Zhong, Yibo
Published: (2025)
by: Zhong, Yibo
Published: (2025)
Similar Items
-
Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
by: Feng, Cheng, et al.
Published: (2026) -
Native Hybrid Attention for Efficient Sequence Modeling
by: Du, Jusen, et al.
Published: (2025) -
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
by: Ren, Yuxin, et al.
Published: (2026) -
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
by: Xie, Yuhan, et al.
Published: (2026) -
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
by: Yang, Shu, et al.
Published: (2025)