RevMUX: Data Multiplexing with Reversible Adapters for Efficient LLM Batch Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yige, Guo, Xu, Zeng, Zhiwei, Miao, Chunyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
by: Xu, Yige, et al.
Published: (2025)
by: Xu, Yige, et al.
Published: (2025)
SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
by: Xu, Yige, et al.
Published: (2025)
by: Xu, Yige, et al.
Published: (2025)
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
by: Wang, Jingchu, et al.
Published: (2026)
by: Wang, Jingchu, et al.
Published: (2026)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
Multi-Bin Batching for Increasing LLM Inference Throughput
by: Guldogan, Ozgur, et al.
Published: (2024)
by: Guldogan, Ozgur, et al.
Published: (2024)
A Closed-Loop Personalized Learning Agent Integrating Neural Cognitive Diagnosis, Bounded-Ability Adaptive Testing, and LLM-Driven Feedback
by: Wang, Zhifeng, et al.
Published: (2025)
by: Wang, Zhifeng, et al.
Published: (2025)
Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
by: Bang, Jihwan, et al.
Published: (2024)
by: Bang, Jihwan, et al.
Published: (2024)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
by: Ebrahimi, Sayna, et al.
Published: (2024)
by: Ebrahimi, Sayna, et al.
Published: (2024)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
READER: Retrieval-Assisted Drafter for Efficient LLM Inference
by: Divilkovskiy, Maxim, et al.
Published: (2025)
by: Divilkovskiy, Maxim, et al.
Published: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
Efficient LLM Inference with Kcache
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
Batch-ICL: Effective, Efficient, and Order-Agnostic In-Context Learning
by: Zhang, Kaiyi, et al.
Published: (2024)
by: Zhang, Kaiyi, et al.
Published: (2024)
Inference-time Alignment in Continuous Space
by: Yuan, Yige, et al.
Published: (2025)
by: Yuan, Yige, et al.
Published: (2025)
Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuning
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
by: Song, Dinghong, et al.
Published: (2025)
by: Song, Dinghong, et al.
Published: (2025)
Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering
by: Dong, Fangan, et al.
Published: (2026)
by: Dong, Fangan, et al.
Published: (2026)
Resolving Word Vagueness with Scenario-guided Adapter for Natural Language Inference
by: Liu, Yonghao, et al.
Published: (2024)
by: Liu, Yonghao, et al.
Published: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning With Adapters
by: Chen, Keyu, et al.
Published: (2024)
by: Chen, Keyu, et al.
Published: (2024)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
by: Agullo, Ferran, et al.
Published: (2025)
by: Agullo, Ferran, et al.
Published: (2025)
ChartAdapter: Large Vision-Language Model for Chart Summarization
by: Xu, Peixin, et al.
Published: (2024)
by: Xu, Peixin, et al.
Published: (2024)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
Adapters Mixup: Mixing Parameter-Efficient Adapters to Enhance the Adversarial Robustness of Fine-tuned Pre-trained Text Classifiers
by: Nguyen, Tuc, et al.
Published: (2024)
by: Nguyen, Tuc, et al.
Published: (2024)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
by: Jiang, Tingyu, et al.
Published: (2025)
by: Jiang, Tingyu, et al.
Published: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
by: Agullo, Ferran, et al.
Published: (2026)
by: Agullo, Ferran, et al.
Published: (2026)
ParaRevSNN: A Parallel Reversible Spiking Neural Network for Efficient Training and Inference
by: Xu, Changqing, et al.
Published: (2025)
by: Xu, Changqing, et al.
Published: (2025)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
Similar Items
-
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
by: Xu, Yige, et al.
Published: (2025) -
SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
by: Xu, Yige, et al.
Published: (2025) -
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024) -
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
by: Xie, Bin, et al.
Published: (2025)