ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ali, Ammar, Mohammad, Baher, Makhov, Denis, Shopkhoev, Dmitriy, Zhussip, Magauiya, Lefkimmiatis, Stamatios |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
von: Makhov, Denis, et al.
Veröffentlicht: (2025)
von: Makhov, Denis, et al.
Veröffentlicht: (2025)
COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
von: Makhov, Denis, et al.
Veröffentlicht: (2026)
von: Makhov, Denis, et al.
Veröffentlicht: (2026)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
von: Mohammad, Baher, et al.
Veröffentlicht: (2025)
von: Mohammad, Baher, et al.
Veröffentlicht: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
A Modular Conditional Diffusion Framework for Image Reconstruction
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2024)
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2024)
SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
GSLoc: Visual Localization with 3D Gaussian Splatting
von: Botashev, Kazii, et al.
Veröffentlicht: (2024)
von: Botashev, Kazii, et al.
Veröffentlicht: (2024)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
CUROCKET: Optimizing ROCKET for GPU
von: Stüven, Ole, et al.
Veröffentlicht: (2026)
von: Stüven, Ole, et al.
Veröffentlicht: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Knapsack Optimization-based Schema Linking for LLM-based Text-to-SQL Generation
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
Defusing Logic Bombs in Symbolic Execution with LLM-Generated Ghost Code
von: Bouras, Dimitrios Stamatios, et al.
Veröffentlicht: (2026)
von: Bouras, Dimitrios Stamatios, et al.
Veröffentlicht: (2026)
Few-Shot Optimized Framework for Hallucination Detection in Resource-Limited NLP Systems
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
Data-Efficient Spectral Classification of Hyperspectral Data Using MiniROCKET and HDC-MiniROCKET
von: Theisen, Nick, et al.
Veröffentlicht: (2025)
von: Theisen, Nick, et al.
Veröffentlicht: (2025)
Private Language Models via Truncated Laplacian Mechanism
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models
von: Basem, Mohamed, et al.
Veröffentlicht: (2024)
von: Basem, Mohamed, et al.
Veröffentlicht: (2024)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
von: Khalid, Haziq Mohammad, et al.
Veröffentlicht: (2025)
von: Khalid, Haziq Mohammad, et al.
Veröffentlicht: (2025)
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
von: Qiao, Ziqing, et al.
Veröffentlicht: (2025)
von: Qiao, Ziqing, et al.
Veröffentlicht: (2025)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
HIT-ROCKET: Hadamard-vector Inner-product Transformer for ROCKET
von: Hao, Wang, et al.
Veröffentlicht: (2025)
von: Hao, Wang, et al.
Veröffentlicht: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
von: Liskavets, Barys, et al.
Veröffentlicht: (2025)
von: Liskavets, Barys, et al.
Veröffentlicht: (2025)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models
von: Amari, Aziz, et al.
Veröffentlicht: (2025)
von: Amari, Aziz, et al.
Veröffentlicht: (2025)
Arabic Large Language Models for Medical Text Generation
von: Allam, Abdulrahman, et al.
Veröffentlicht: (2025)
von: Allam, Abdulrahman, et al.
Veröffentlicht: (2025)
Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
Enhancing and Accelerating Large Language Models via Instruction-Aware Contextual Compression
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models
von: Zhao, Runsong, et al.
Veröffentlicht: (2024)
von: Zhao, Runsong, et al.
Veröffentlicht: (2024)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
von: Wang, Weilan, et al.
Veröffentlicht: (2025)
von: Wang, Weilan, et al.
Veröffentlicht: (2025)
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
Fewer Truncations Improve Language Modeling
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization
von: Hosseinbeigi, Sara Bourbour, et al.
Veröffentlicht: (2025)
von: Hosseinbeigi, Sara Bourbour, et al.
Veröffentlicht: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
Ranked List Truncation for Large Language Model-based Re-Ranking
von: Meng, Chuan, et al.
Veröffentlicht: (2024)
von: Meng, Chuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
von: Makhov, Denis, et al.
Veröffentlicht: (2025) -
COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
von: Makhov, Denis, et al.
Veröffentlicht: (2026) -
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025) -
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
von: Mohammad, Baher, et al.
Veröffentlicht: (2025) -
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)