LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xuan, Zhang, Fengzhuo, Du, Cunxiao, Du, Chao, Pang, Tianyu, Gao, Wei, Lin, Min |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
par: Yang, Penghui, et autres
Publié: (2025)
par: Yang, Penghui, et autres
Publié: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
par: Gu, Xiangming, et autres
Publié: (2024)
par: Gu, Xiangming, et autres
Publié: (2024)
Demystifying the Slash Pattern in Attention: The Role of RoPE
par: Cheng, Yuan, et autres
Publié: (2026)
par: Cheng, Yuan, et autres
Publié: (2026)
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
par: Wang, Haonan, et autres
Publié: (2024)
par: Wang, Haonan, et autres
Publié: (2024)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
par: Zhang, Xuan, et autres
Publié: (2025)
par: Zhang, Xuan, et autres
Publié: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
par: Zhang, Xuan, et autres
Publié: (2024)
par: Zhang, Xuan, et autres
Publié: (2024)
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One
par: Li, Tianlin, et autres
Publié: (2024)
par: Li, Tianlin, et autres
Publié: (2024)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
par: Min, Rui, et autres
Publié: (2025)
par: Min, Rui, et autres
Publié: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
par: Li, Xingyao, et autres
Publié: (2026)
par: Li, Xingyao, et autres
Publié: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
par: Hou, Yunlong, et autres
Publié: (2025)
par: Hou, Yunlong, et autres
Publié: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
par: Xia, Heming, et autres
Publié: (2024)
par: Xia, Heming, et autres
Publié: (2024)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
par: Zheng, Xiaosen, et autres
Publié: (2024)
par: Zheng, Xiaosen, et autres
Publié: (2024)
A Closer Look at Machine Unlearning for Large Language Models
par: Yuan, Xiaojian, et autres
Publié: (2024)
par: Yuan, Xiaojian, et autres
Publié: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
par: Gao, Kuofeng, et autres
Publié: (2024)
par: Gao, Kuofeng, et autres
Publié: (2024)
Muon Outperforms Adam in Tail-End Associative Memory Learning
par: Wang, Shuche, et autres
Publié: (2025)
par: Wang, Shuche, et autres
Publié: (2025)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
par: Xia, Heming, et autres
Publié: (2025)
par: Xia, Heming, et autres
Publié: (2025)
Flora: Effortless Context Construction to Arbitrary Length and Scale
par: Chen, Tianxiang, et autres
Publié: (2025)
par: Chen, Tianxiang, et autres
Publié: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
par: Gao, Hongcheng, et autres
Publié: (2024)
par: Gao, Hongcheng, et autres
Publié: (2024)
Benchmarking Large Multimodal Models against Common Corruptions
par: Zhang, Jiawei, et autres
Publié: (2024)
par: Zhang, Jiawei, et autres
Publié: (2024)
Long-Context Language Modeling with Parallel Context Encoding
par: Yen, Howard, et autres
Publié: (2024)
par: Yen, Howard, et autres
Publié: (2024)
Revisiting the Markov Property for Machine Translation
par: Du, Cunxiao, et autres
Publié: (2024)
par: Du, Cunxiao, et autres
Publié: (2024)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
par: Zheng, Xiaosen, et autres
Publié: (2024)
par: Zheng, Xiaosen, et autres
Publié: (2024)
Scalable Token-Level Hallucination Detection in Large Language Models
par: Min, Rui, et autres
Publié: (2026)
par: Min, Rui, et autres
Publié: (2026)
Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
par: Wang, Haonan, et autres
Publié: (2025)
par: Wang, Haonan, et autres
Publié: (2025)
LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
par: Huang, Chengsong, et autres
Publié: (2023)
par: Huang, Chengsong, et autres
Publié: (2023)
Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
par: Du, Hanwen, et autres
Publié: (2025)
par: Du, Hanwen, et autres
Publié: (2025)
Test-Time Backdoor Attacks on Multimodal Large Language Models
par: Lu, Dong, et autres
Publié: (2024)
par: Lu, Dong, et autres
Publié: (2024)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
par: Fang, Hao, et autres
Publié: (2025)
par: Fang, Hao, et autres
Publié: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
par: Ai, Xuan, et autres
Publié: (2026)
par: Ai, Xuan, et autres
Publié: (2026)
Purifying Large Language Models by Ensembling a Small Language Model
par: Li, Tianlin, et autres
Publié: (2024)
par: Li, Tianlin, et autres
Publié: (2024)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
par: Luo, Renjie, et autres
Publié: (2025)
par: Luo, Renjie, et autres
Publié: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
par: Lin, Gang, et autres
Publié: (2026)
par: Lin, Gang, et autres
Publié: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
par: Lu, Yi, et autres
Publié: (2024)
par: Lu, Yi, et autres
Publié: (2024)
From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems
par: He, Jianliang, et autres
Publié: (2024)
par: He, Jianliang, et autres
Publié: (2024)
Lifelong Safety Alignment for Language Models
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
par: Li, Ziyue, et autres
Publié: (2024)
par: Li, Ziyue, et autres
Publié: (2024)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
par: Ye, Xi, et autres
Publié: (2025)
par: Ye, Xi, et autres
Publié: (2025)
Efficiently Computing Susceptibility to Context in Language Models
par: Liu, Tianyu, et autres
Publié: (2024)
par: Liu, Tianyu, et autres
Publié: (2024)
Bootstrapping Language Models with DPO Implicit Rewards
par: Chen, Changyu, et autres
Publié: (2024)
par: Chen, Changyu, et autres
Publié: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
par: Gu, Zhuohan, et autres
Publié: (2024)
par: Gu, Zhuohan, et autres
Publié: (2024)
Documents similaires
-
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
par: Yang, Penghui, et autres
Publié: (2025) -
When Attention Sink Emerges in Language Models: An Empirical View
par: Gu, Xiangming, et autres
Publié: (2024) -
Demystifying the Slash Pattern in Attention: The Role of RoPE
par: Cheng, Yuan, et autres
Publié: (2026) -
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
par: Wang, Haonan, et autres
Publié: (2024) -
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
par: Zhang, Xuan, et autres
Publié: (2025)