4bit-Quantization in Vector-Embedding for RAG
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Jeong, Taehee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightweight Relevance Grader in RAG
von: Jeong, Taehee
Veröffentlicht: (2025)
von: Jeong, Taehee
Veröffentlicht: (2025)
ICQuant: Index Coding enables Low-bit LLM Quantization
von: Li, Xinlin, et al.
Veröffentlicht: (2025)
von: Li, Xinlin, et al.
Veröffentlicht: (2025)
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
Representation Collapsing Problems in Vector Quantization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
von: Zhai, Linwei, et al.
Veröffentlicht: (2026)
von: Zhai, Linwei, et al.
Veröffentlicht: (2026)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
von: Zhang, Michael S., et al.
Veröffentlicht: (2025)
von: Zhang, Michael S., et al.
Veröffentlicht: (2025)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
von: Xu, Chen, et al.
Veröffentlicht: (2025)
von: Xu, Chen, et al.
Veröffentlicht: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
4-bit Shampoo for Memory-Efficient Network Training
von: Wang, Sike, et al.
Veröffentlicht: (2024)
von: Wang, Sike, et al.
Veröffentlicht: (2024)
Vector Quantization in the Brain: Grid-like Codes in World Models
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
Educational Cone Model in Embedding Vector Spaces
von: Ehara, Yo
Veröffentlicht: (2025)
von: Ehara, Yo
Veröffentlicht: (2025)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
von: Song, Yurun, et al.
Veröffentlicht: (2025)
von: Song, Yurun, et al.
Veröffentlicht: (2025)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
von: Lee, Geonho, et al.
Veröffentlicht: (2024)
von: Lee, Geonho, et al.
Veröffentlicht: (2024)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
Divide et Calibra: Multiclass Local Calibration via Vector Quantization
von: Barbera, Cesare, et al.
Veröffentlicht: (2026)
von: Barbera, Cesare, et al.
Veröffentlicht: (2026)
VQ-SAD: Vector Quantized Structure Aware Diffusion For Molecule Generation
von: Noravesh, Farshad, et al.
Veröffentlicht: (2026)
von: Noravesh, Farshad, et al.
Veröffentlicht: (2026)
Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization
von: Luu, Tung M., et al.
Veröffentlicht: (2024)
von: Luu, Tung M., et al.
Veröffentlicht: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning
von: Zhai, Zian, et al.
Veröffentlicht: (2025)
von: Zhai, Zian, et al.
Veröffentlicht: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
SDQ: Sparse Decomposed Quantization for LLM Inference
von: Jeong, Geonhwa, et al.
Veröffentlicht: (2024)
von: Jeong, Geonhwa, et al.
Veröffentlicht: (2024)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
von: Yue, Yuxuan, et al.
Veröffentlicht: (2025)
von: Yue, Yuxuan, et al.
Veröffentlicht: (2025)
SmaAT-QMix-UNet: A Parameter-Efficient Vector-Quantized UNet for Precipitation Nowcasting
von: Stavrou, Nikolas, et al.
Veröffentlicht: (2026)
von: Stavrou, Nikolas, et al.
Veröffentlicht: (2026)
TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding
von: Zhang, Haochuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haochuan, et al.
Veröffentlicht: (2025)
FoldToken: Learning Protein Language via Vector Quantization and Beyond
von: Gao, Zhangyang, et al.
Veröffentlicht: (2024)
von: Gao, Zhangyang, et al.
Veröffentlicht: (2024)
BeamVQ: Beam Search with Vector Quantization to Mitigate Data Scarcity in Physical Spatiotemporal Forecasting
von: Wang, Weiyan, et al.
Veröffentlicht: (2025)
von: Wang, Weiyan, et al.
Veröffentlicht: (2025)
Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
von: Morita, Takashi
Veröffentlicht: (2025)
von: Morita, Takashi
Veröffentlicht: (2025)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
von: Zhao, Yanjun, et al.
Veröffentlicht: (2024)
von: Zhao, Yanjun, et al.
Veröffentlicht: (2024)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
von: Xia, Junhao, et al.
Veröffentlicht: (2025)
von: Xia, Junhao, et al.
Veröffentlicht: (2025)
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
von: Bhatia, Randeep, et al.
Veröffentlicht: (2025)
von: Bhatia, Randeep, et al.
Veröffentlicht: (2025)
Matmul or No Matmul in the Era of 1-bit LLMs
von: Malekar, Jinendra, et al.
Veröffentlicht: (2024)
von: Malekar, Jinendra, et al.
Veröffentlicht: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
VQSynery: Robust Drug Synergy Prediction With Vector Quantization Mechanism
von: Wu, Jiawei, et al.
Veröffentlicht: (2024)
von: Wu, Jiawei, et al.
Veröffentlicht: (2024)
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
von: Ling, Tianheng, et al.
Veröffentlicht: (2024)
von: Ling, Tianheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Lightweight Relevance Grader in RAG
von: Jeong, Taehee
Veröffentlicht: (2025) -
ICQuant: Index Coding enables Low-bit LLM Quantization
von: Li, Xinlin, et al.
Veröffentlicht: (2025) -
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025) -
Representation Collapsing Problems in Vector Quantization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)