Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Peijie, Tang, Zhenheng, Liu, Xiang, Li, Lujun, Chu, Xiaowen, Li, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
por: Tang, Zhenheng, et al.
Publicado: (2025)
por: Tang, Zhenheng, et al.
Publicado: (2025)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
por: Dong, Peijie, et al.
Publicado: (2024)
por: Dong, Peijie, et al.
Publicado: (2024)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
por: Lai, Kunfeng, et al.
Publicado: (2025)
por: Lai, Kunfeng, et al.
Publicado: (2025)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
por: Dong, Peijie, et al.
Publicado: (2024)
por: Dong, Peijie, et al.
Publicado: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
por: Li, Qi, et al.
Publicado: (2025)
por: Li, Qi, et al.
Publicado: (2025)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
por: Tang, Zichen, et al.
Publicado: (2024)
por: Tang, Zichen, et al.
Publicado: (2024)
Delta Decompression for MoE-based LLMs Compression
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
LPZero: Language Model Zero-cost Proxy Search from Zero
por: Dong, Peijie, et al.
Publicado: (2024)
por: Dong, Peijie, et al.
Publicado: (2024)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
por: Tang, Zhenheng, et al.
Publicado: (2024)
por: Tang, Zhenheng, et al.
Publicado: (2024)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
por: Dong, Peijie, et al.
Publicado: (2026)
por: Dong, Peijie, et al.
Publicado: (2026)
FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion
por: Tang, Zhenheng, et al.
Publicado: (2024)
por: Tang, Zhenheng, et al.
Publicado: (2024)
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
por: Liu, Xiang, et al.
Publicado: (2026)
por: Liu, Xiang, et al.
Publicado: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
Rethinking Deep Research from the Perspective of Web Content Distribution Matching
por: Yu, Zixuan, et al.
Publicado: (2026)
por: Yu, Zixuan, et al.
Publicado: (2026)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
por: Zhang, Te, et al.
Publicado: (2025)
por: Zhang, Te, et al.
Publicado: (2025)
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
por: Pan, Xinglin, et al.
Publicado: (2025)
por: Pan, Xinglin, et al.
Publicado: (2025)
On the Spectral Flattening of Quantized Embeddings
por: Huang, Junlin, et al.
Publicado: (2026)
por: Huang, Junlin, et al.
Publicado: (2026)
AQUATIC-Diff: Additive Quantization for Truly Tiny Compressed Diffusion Models
por: Hasan, Adil, et al.
Publicado: (2025)
por: Hasan, Adil, et al.
Publicado: (2025)
CompAct: Compressed Activations for Memory-Efficient LLM Training
por: Shamshoum, Yara, et al.
Publicado: (2024)
por: Shamshoum, Yara, et al.
Publicado: (2024)
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
por: Li, Qi, et al.
Publicado: (2024)
por: Li, Qi, et al.
Publicado: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
por: Queipo-de-Llano, Enrique, et al.
Publicado: (2025)
por: Queipo-de-Llano, Enrique, et al.
Publicado: (2025)
FedImpro: Measuring and Improving Client Update in Federated Learning
por: Tang, Zhenheng, et al.
Publicado: (2024)
por: Tang, Zhenheng, et al.
Publicado: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
por: Fan, Ruibo, et al.
Publicado: (2026)
por: Fan, Ruibo, et al.
Publicado: (2026)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
por: Maveli, Nickil, et al.
Publicado: (2026)
por: Maveli, Nickil, et al.
Publicado: (2026)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
por: Xiang, Maoyang, et al.
Publicado: (2025)
por: Xiang, Maoyang, et al.
Publicado: (2025)
MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
por: Li, Guangyan, et al.
Publicado: (2025)
por: Li, Guangyan, et al.
Publicado: (2025)
RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
por: He, Xin, et al.
Publicado: (2025)
por: He, Xin, et al.
Publicado: (2025)
VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting
por: Tang, Yujin, et al.
Publicado: (2024)
por: Tang, Yujin, et al.
Publicado: (2024)
Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm
por: Wei, Wen-Da, et al.
Publicado: (2026)
por: Wei, Wen-Da, et al.
Publicado: (2026)
Skill Reuse as Compression in Agentic RL
por: Xu, Zhikun, et al.
Publicado: (2026)
por: Xu, Zhikun, et al.
Publicado: (2026)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
por: Tang, Bohan, et al.
Publicado: (2025)
por: Tang, Bohan, et al.
Publicado: (2025)
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
por: Wu, Mingyuan, et al.
Publicado: (2025)
por: Wu, Mingyuan, et al.
Publicado: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
por: Yu, Xiaoming, et al.
Publicado: (2026)
por: Yu, Xiaoming, et al.
Publicado: (2026)
Compressed Empirical Measures (in finite dimensions)
por: Grünewälder, Steffen
Publicado: (2022)
por: Grünewälder, Steffen
Publicado: (2022)
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
por: Li, Wenshuo, et al.
Publicado: (2024)
por: Li, Wenshuo, et al.
Publicado: (2024)
Elastic-dLLM: Position Preserving Context Compression and Augmentation of Diffusion LLMs
por: Wu, Junyi, et al.
Publicado: (2026)
por: Wu, Junyi, et al.
Publicado: (2026)
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
por: Tang, Zichen, et al.
Publicado: (2026)
por: Tang, Zichen, et al.
Publicado: (2026)
Ejemplares similares
-
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
por: Tang, Zhenheng, et al.
Publicado: (2025) -
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
por: Dong, Peijie, et al.
Publicado: (2024) -
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
por: Lai, Kunfeng, et al.
Publicado: (2025) -
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
por: Dong, Peijie, et al.
Publicado: (2024) -
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)