Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Wanlong, Xiao, Yichen, Zeng, Dingyi, Zhao, Hongyang, Chen, Wenyu, Zhang, Malu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Utilizing Contextual Clues and Role Correlations for Enhancing Document-level Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2023)
von: Liu, Wanlong, et al.
Veröffentlicht: (2023)
Beyond Single-Event Extraction: Towards Efficient Document-Level Multi-Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
A Compressive Memory-based Retrieval Approach for Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
MLPs Compass: What is learned when MLPs are combined with PLMs?
von: Zhou, Li, et al.
Veröffentlicht: (2024)
von: Zhou, Li, et al.
Veröffentlicht: (2024)
DPGNN: Dual-Perception Graph Neural Network for Representation Learning
von: Zhou, Li, et al.
Veröffentlicht: (2021)
von: Zhou, Li, et al.
Veröffentlicht: (2021)
Channel-Wise Mixed-Precision Quantization for Large Language Models
von: Chen, Zihan, et al.
Veröffentlicht: (2024)
von: Chen, Zihan, et al.
Veröffentlicht: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
Give Us the Facts: Enhancing Large Language Models with Knowledge Graphs for Fact-aware Language Modeling
von: Yang, Linyao, et al.
Veröffentlicht: (2023)
von: Yang, Linyao, et al.
Veröffentlicht: (2023)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
OneBit: Towards Extremely Low-bit Large Language Models
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
von: Nanda, Amitash, et al.
Veröffentlicht: (2024)
von: Nanda, Amitash, et al.
Veröffentlicht: (2024)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
Multi-Bit Distortion-Free Watermarking for Large Language Models
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models
von: Guo, Wenbin, et al.
Veröffentlicht: (2025)
von: Guo, Wenbin, et al.
Veröffentlicht: (2025)
Two Heads Are Better Than One: Integrating Knowledge from Knowledge Graphs and Large Language Models for Entity Alignment
von: Yang, Linyao, et al.
Veröffentlicht: (2024)
von: Yang, Linyao, et al.
Veröffentlicht: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
von: Ping, Bowen, et al.
Veröffentlicht: (2024)
von: Ping, Bowen, et al.
Veröffentlicht: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
How Creative Are Large Language Models in Generating Molecules?
von: Tao, Wen, et al.
Veröffentlicht: (2026)
von: Tao, Wen, et al.
Veröffentlicht: (2026)
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models for Code-Mixed Data Augmentation in Sentiment Analysis
von: Zeng, Linda
Veröffentlicht: (2024)
von: Zeng, Linda
Veröffentlicht: (2024)
Ähnliche Einträge
-
Utilizing Contextual Clues and Role Correlations for Enhancing Document-level Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2023) -
Beyond Single-Event Extraction: Towards Efficient Document-Level Multi-Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2024) -
A Compressive Memory-based Retrieval Approach for Event Argument Extraction
von: Liu, Wanlong, et al.
Veröffentlicht: (2024) -
MLPs Compass: What is learned when MLPs are combined with PLMs?
von: Zhou, Li, et al.
Veröffentlicht: (2024) -
DPGNN: Dual-Perception Graph Neural Network for Representation Learning
von: Zhou, Li, et al.
Veröffentlicht: (2021)