4bit-Quantization in Vector-Embedding for RAG
Fuente:
arXiv
Salvato in:
| Autore principale: | Jeong, Taehee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lightweight Relevance Grader in RAG
di: Jeong, Taehee
Pubblicazione: (2025)
di: Jeong, Taehee
Pubblicazione: (2025)
ICQuant: Index Coding enables Low-bit LLM Quantization
di: Li, Xinlin, et al.
Pubblicazione: (2025)
di: Li, Xinlin, et al.
Pubblicazione: (2025)
Low-bit Model Quantization for Deep Neural Networks: A Survey
di: Liu, Kai, et al.
Pubblicazione: (2025)
di: Liu, Kai, et al.
Pubblicazione: (2025)
Representation Collapsing Problems in Vector Quantization
di: Zhao, Wenhao, et al.
Pubblicazione: (2024)
di: Zhao, Wenhao, et al.
Pubblicazione: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
di: Nielsen, Jacob, et al.
Pubblicazione: (2025)
di: Nielsen, Jacob, et al.
Pubblicazione: (2025)
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
di: Zhai, Linwei, et al.
Pubblicazione: (2026)
di: Zhai, Linwei, et al.
Pubblicazione: (2026)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
di: Zhang, Michael S., et al.
Pubblicazione: (2025)
di: Zhang, Michael S., et al.
Pubblicazione: (2025)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
di: Xu, Chen, et al.
Pubblicazione: (2025)
di: Xu, Chen, et al.
Pubblicazione: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
di: Xia, Haojun, et al.
Pubblicazione: (2025)
di: Xia, Haojun, et al.
Pubblicazione: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2025)
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2025)
4-bit Shampoo for Memory-Efficient Network Training
di: Wang, Sike, et al.
Pubblicazione: (2024)
di: Wang, Sike, et al.
Pubblicazione: (2024)
Vector Quantization in the Brain: Grid-like Codes in World Models
di: Peng, Xiangyuan, et al.
Pubblicazione: (2025)
di: Peng, Xiangyuan, et al.
Pubblicazione: (2025)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
Educational Cone Model in Embedding Vector Spaces
di: Ehara, Yo
Pubblicazione: (2025)
di: Ehara, Yo
Pubblicazione: (2025)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
di: Song, Yurun, et al.
Pubblicazione: (2025)
di: Song, Yurun, et al.
Pubblicazione: (2025)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
di: Lee, Geonho, et al.
Pubblicazione: (2024)
di: Lee, Geonho, et al.
Pubblicazione: (2024)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
di: Zhang, Xi, et al.
Pubblicazione: (2025)
di: Zhang, Xi, et al.
Pubblicazione: (2025)
Divide et Calibra: Multiclass Local Calibration via Vector Quantization
di: Barbera, Cesare, et al.
Pubblicazione: (2026)
di: Barbera, Cesare, et al.
Pubblicazione: (2026)
VQ-SAD: Vector Quantized Structure Aware Diffusion For Molecule Generation
di: Noravesh, Farshad, et al.
Pubblicazione: (2026)
di: Noravesh, Farshad, et al.
Pubblicazione: (2026)
Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization
di: Luu, Tung M., et al.
Pubblicazione: (2024)
di: Luu, Tung M., et al.
Pubblicazione: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning
di: Zhai, Zian, et al.
Pubblicazione: (2025)
di: Zhai, Zian, et al.
Pubblicazione: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
SDQ: Sparse Decomposed Quantization for LLM Inference
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
di: Yue, Yuxuan, et al.
Pubblicazione: (2025)
di: Yue, Yuxuan, et al.
Pubblicazione: (2025)
SmaAT-QMix-UNet: A Parameter-Efficient Vector-Quantized UNet for Precipitation Nowcasting
di: Stavrou, Nikolas, et al.
Pubblicazione: (2026)
di: Stavrou, Nikolas, et al.
Pubblicazione: (2026)
TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding
di: Zhang, Haochuan, et al.
Pubblicazione: (2025)
di: Zhang, Haochuan, et al.
Pubblicazione: (2025)
FoldToken: Learning Protein Language via Vector Quantization and Beyond
di: Gao, Zhangyang, et al.
Pubblicazione: (2024)
di: Gao, Zhangyang, et al.
Pubblicazione: (2024)
BeamVQ: Beam Search with Vector Quantization to Mitigate Data Scarcity in Physical Spatiotemporal Forecasting
di: Wang, Weiyan, et al.
Pubblicazione: (2025)
di: Wang, Weiyan, et al.
Pubblicazione: (2025)
Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
di: Morita, Takashi
Pubblicazione: (2025)
di: Morita, Takashi
Pubblicazione: (2025)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
di: Xia, Junhao, et al.
Pubblicazione: (2025)
di: Xia, Junhao, et al.
Pubblicazione: (2025)
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
di: Bhatia, Randeep, et al.
Pubblicazione: (2025)
di: Bhatia, Randeep, et al.
Pubblicazione: (2025)
Matmul or No Matmul in the Era of 1-bit LLMs
di: Malekar, Jinendra, et al.
Pubblicazione: (2024)
di: Malekar, Jinendra, et al.
Pubblicazione: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
di: Son, Donghyun, et al.
Pubblicazione: (2025)
di: Son, Donghyun, et al.
Pubblicazione: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
di: Zhong, Zhengjia, et al.
Pubblicazione: (2026)
di: Zhong, Zhengjia, et al.
Pubblicazione: (2026)
VQSynery: Robust Drug Synergy Prediction With Vector Quantization Mechanism
di: Wu, Jiawei, et al.
Pubblicazione: (2024)
di: Wu, Jiawei, et al.
Pubblicazione: (2024)
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
di: Ling, Tianheng, et al.
Pubblicazione: (2024)
di: Ling, Tianheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Lightweight Relevance Grader in RAG
di: Jeong, Taehee
Pubblicazione: (2025) -
ICQuant: Index Coding enables Low-bit LLM Quantization
di: Li, Xinlin, et al.
Pubblicazione: (2025) -
Low-bit Model Quantization for Deep Neural Networks: A Survey
di: Liu, Kai, et al.
Pubblicazione: (2025) -
Representation Collapsing Problems in Vector Quantization
di: Zhao, Wenhao, et al.
Pubblicazione: (2024) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)