RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sung, Yi-Lin, Yadav, Prateek, Li, Jialu, Yoon, Jaehong, Bansal, Mohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
Glider: Global and Local Instruction-Driven Expert Router
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
von: Jiang, Yichen, et al.
Veröffentlicht: (2024)
von: Jiang, Yichen, et al.
Veröffentlicht: (2024)
What Matters for Model Merging at Scale?
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
von: Maharana, Adyasha, et al.
Veröffentlicht: (2024)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2024)
How Important Is Tokenization in French Medical Masked Language Models?
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
Learning to Route LLMs with Confidence Tokens
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
Does Biomedical Training Lead to Better Medical Performance?
von: Dada, Amin, et al.
Veröffentlicht: (2024)
von: Dada, Amin, et al.
Veröffentlicht: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
When Does Multimodality Lead to Better Time Series Forecasting?
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2025)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
Soft Self-Consistency Improves Language Model Agents
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
Interpreting the Effects of Quantization on LLMs
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
von: Xiao, Hanqi, et al.
Veröffentlicht: (2026)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2026)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023) -
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023) -
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024) -
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)