LPCD: Unified Framework from Layer-Wise to Submodule Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ichikawa, Yuma, Fujimoto, Yudai, Sakai, Akira |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
Sign Lock-In: Randomly Initialized Weight Signs Persist and Bottleneck Sub-Bit Model Compression
von: Sakai, Akira, et al.
Veröffentlicht: (2026)
von: Sakai, Akira, et al.
Veröffentlicht: (2026)
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
von: Fujisawa, Yoshihiko, et al.
Veröffentlicht: (2026)
von: Fujisawa, Yoshihiko, et al.
Veröffentlicht: (2026)
PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
OneComp: One-Line Revolution for Generative AI Model Compression
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2026)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2026)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
von: Arai, Yamato, et al.
Veröffentlicht: (2025)
von: Arai, Yamato, et al.
Veröffentlicht: (2025)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
von: Sim, Woo Seob, et al.
Veröffentlicht: (2026)
von: Sim, Woo Seob, et al.
Veröffentlicht: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
von: Ye, Donald
Veröffentlicht: (2026)
von: Ye, Donald
Veröffentlicht: (2026)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
von: Kato, Mariko, et al.
Veröffentlicht: (2025)
von: Kato, Mariko, et al.
Veröffentlicht: (2025)
Token-based Decision Criteria Are Suboptimal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
von: Jiang, Jiachen, et al.
Veröffentlicht: (2024)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2024)
EVE-Agent: Evidence-Verifiable Self-Evolving Agents
von: Arai, Yamato, et al.
Veröffentlicht: (2026)
von: Arai, Yamato, et al.
Veröffentlicht: (2026)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
A Unified Framework for Model Editing
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain
von: Okochi, Yuma, et al.
Veröffentlicht: (2026)
von: Okochi, Yuma, et al.
Veröffentlicht: (2026)
Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2026)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2026)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024)
von: He, Junhui, et al.
Veröffentlicht: (2024)
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
von: Zhang, Jinhao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinhao, et al.
Veröffentlicht: (2025)
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
von: Okutomi, Akira
Veröffentlicht: (2025)
von: Okutomi, Akira
Veröffentlicht: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
von: Cunegatti, Elia, et al.
Veröffentlicht: (2026)
von: Cunegatti, Elia, et al.
Veröffentlicht: (2026)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
von: Yang, Daniel, et al.
Veröffentlicht: (2026)
von: Yang, Daniel, et al.
Veröffentlicht: (2026)
Interpreting the Effects of Quantization on LLMs
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
Combating Confirmation Bias: A Unified Pseudo-Labeling Framework for Entity Alignment
von: Ding, Qijie, et al.
Veröffentlicht: (2023)
von: Ding, Qijie, et al.
Veröffentlicht: (2023)
uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
von: Zou, Heming, et al.
Veröffentlicht: (2025)
von: Zou, Heming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025) -
Sign Lock-In: Randomly Initialized Weight Signs Persist and Bottleneck Sub-Bit Model Compression
von: Sakai, Akira, et al.
Veröffentlicht: (2026) -
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
von: Fujisawa, Yoshihiko, et al.
Veröffentlicht: (2026) -
PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025) -
OneComp: One-Line Revolution for Generative AI Model Compression
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2026)