Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xiao, Chaojun, Luo, Yuqi, Zhang, Wenbin, Zhang, Pengle, Han, Xu, Lin, Yankai, Zhang, Zhengyan, Xie, Ruobing, Liu, Zhiyuan, Sun, Maosong, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential Recommendation
par: Qu, Zekai, et autres
Publié: (2024)
par: Qu, Zekai, et autres
Publié: (2024)
Representation Learning for Natural Language Processing
par: Liu, Zhiyuan, et autres
Publié: (2020)
par: Liu, Zhiyuan, et autres
Publié: (2020)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
par: Huang, Yuxiang, et autres
Publié: (2025)
par: Huang, Yuxiang, et autres
Publié: (2025)
Robust and Scalable Model Editing for Large Language Models
par: Chen, Yingfa, et autres
Publié: (2024)
par: Chen, Yingfa, et autres
Publié: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
par: Zhang, Jintao, et autres
Publié: (2024)
par: Zhang, Jintao, et autres
Publié: (2024)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation
par: Wu, Zhichao, et autres
Publié: (2026)
par: Wu, Zhichao, et autres
Publié: (2026)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
par: Song, Chenyang, et autres
Publié: (2025)
par: Song, Chenyang, et autres
Publié: (2025)
LogLite: Lightweight Plug-and-Play Streaming Log Compression
par: Tang, Benzhao, et autres
Publié: (2025)
par: Tang, Benzhao, et autres
Publié: (2025)
ID-centric Pre-training for Recommendation
par: Wu, Yiqing, et autres
Publié: (2024)
par: Wu, Yiqing, et autres
Publié: (2024)
CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices
par: Zhao, Weilin, et autres
Publié: (2023)
par: Zhao, Weilin, et autres
Publié: (2023)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
par: Luo, Yuqi, et autres
Publié: (2024)
par: Luo, Yuqi, et autres
Publié: (2024)
GS-Net: Generalizable Plug-and-Play 3D Gaussian Splatting Module
par: Zhang, Yichen, et autres
Publié: (2024)
par: Zhang, Yichen, et autres
Publié: (2024)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
par: Guo, Yiju, et autres
Publié: (2024)
par: Guo, Yiju, et autres
Publié: (2024)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
par: Gao, Cheng, et autres
Publié: (2025)
par: Gao, Cheng, et autres
Publié: (2025)
Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
par: Gao, Cheng, et autres
Publié: (2024)
par: Gao, Cheng, et autres
Publié: (2024)
Curriculum-scheduled Knowledge Distillation from Multiple Pre-trained Teachers for Multi-domain Sequential Recommendation
par: Sun, Wenqi, et autres
Publié: (2024)
par: Sun, Wenqi, et autres
Publié: (2024)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
par: Cai, Chang, et autres
Publié: (2025)
par: Cai, Chang, et autres
Publié: (2025)
Plug-in Diffusion Model for Sequential Recommendation
par: Ma, Haokai, et autres
Publié: (2024)
par: Ma, Haokai, et autres
Publié: (2024)
Decoupled Alignment for Robust Plug-and-Play Adaptation
par: Luo, Haozheng, et autres
Publié: (2024)
par: Luo, Haozheng, et autres
Publié: (2024)
TGPP: Trajectory-Guided Plug-and-Play Priors for Sparse Radio Map Reconstruction
par: Zhang, Jiawen, et autres
Publié: (2026)
par: Zhang, Jiawen, et autres
Publié: (2026)
Densing Law of LLMs
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Image Recovery
par: Cai, Chang, et autres
Publié: (2025)
par: Cai, Chang, et autres
Publié: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
Plug-and-Play Versatile Compressed Video Enhancement
par: Zeng, Huimin, et autres
Publié: (2025)
par: Zeng, Huimin, et autres
Publié: (2025)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
par: Zhao, Weilin, et autres
Publié: (2024)
par: Zhao, Weilin, et autres
Publié: (2024)
Accelerating Cross‐Scenario Metasurface Adaptability with Plug‐and‐Play Kernel
par: Nanxuan Wu, et autres
Publié: (2025)
par: Nanxuan Wu, et autres
Publié: (2025)
PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
par: Liu, Jinhan, et autres
Publié: (2026)
par: Liu, Jinhan, et autres
Publié: (2026)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
par: Cai, Xin, et autres
Publié: (2026)
par: Cai, Xin, et autres
Publié: (2026)
Lifelike Agility and Play in Quadrupedal Robots using Reinforcement Learning and Generative Pre-trained Models
par: Han, Lei, et autres
Publié: (2023)
par: Han, Lei, et autres
Publié: (2023)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
par: Li, Miaoran, et autres
Publié: (2023)
par: Li, Miaoran, et autres
Publié: (2023)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
par: Cui, Ganqu, et autres
Publié: (2023)
par: Cui, Ganqu, et autres
Publié: (2023)
From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
par: Ma, Changhai, et autres
Publié: (2025)
par: Ma, Changhai, et autres
Publié: (2025)
UTCS: Effective Unsupervised Temporal Community Search with Pre-training of Temporal Dynamics and Subgraph Knowledge
par: Zhang, Yue, et autres
Publié: (2025)
par: Zhang, Yue, et autres
Publié: (2025)
Plug-and-Play Transformer Modules for Test-Time Adaptation
par: Chang, Xiangyu, et autres
Publié: (2024)
par: Chang, Xiangyu, et autres
Publié: (2024)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
par: Wang, Xing, et autres
Publié: (2025)
par: Wang, Xing, et autres
Publié: (2025)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
par: Xiong, Minhao, et autres
Publié: (2025)
par: Xiong, Minhao, et autres
Publié: (2025)
Documents similaires
-
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024) -
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
par: Xiao, Chaojun, et autres
Publié: (2024) -
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
par: Zhang, Zhengyan, et autres
Publié: (2024) -
The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential Recommendation
par: Qu, Zekai, et autres
Publié: (2024) -
Representation Learning for Natural Language Processing
par: Liu, Zhiyuan, et autres
Publié: (2020)