Ascend-RaBitQ: Heterogeneous NPU-CPU Acceleration of Billion-Scale Similarity Search with 1-bit Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Fujun, Ye, Chuyue, Cai, Huaxiang, Lv, Zetao, Cui, Baolong, Yan, Wenru, Zhan, Chao, Zhang, Zigang, Yi, Hao, Xiang, Jie, Li, Xiabing, Gai, Yuhang, Zhang, Ziyang, Zheng, Pengfei, Du, Yunfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
von: Gao, Jianyang, et al.
Veröffentlicht: (2024)
von: Gao, Jianyang, et al.
Veröffentlicht: (2024)
Revisiting RaBitQ and TurboQuant: A Symmetric Comparison of Methods, Theory, and Experiments
von: Gao, Jianyang, et al.
Veröffentlicht: (2026)
von: Gao, Jianyang, et al.
Veröffentlicht: (2026)
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
von: Shi, Jifan, et al.
Veröffentlicht: (2026)
von: Shi, Jifan, et al.
Veröffentlicht: (2026)
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
von: Wu, Jiehao, et al.
Veröffentlicht: (2026)
von: Wu, Jiehao, et al.
Veröffentlicht: (2026)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2026)
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2026)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
von: Bouquet, Yann, et al.
Veröffentlicht: (2026)
von: Bouquet, Yann, et al.
Veröffentlicht: (2026)
WindVE: Collaborative CPU-NPU Vector Embedding
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
von: Li, Yuhang, et al.
Veröffentlicht: (2024)
von: Li, Yuhang, et al.
Veröffentlicht: (2024)
Ascend-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads
von: Dhar, Aritra, et al.
Veröffentlicht: (2024)
von: Dhar, Aritra, et al.
Veröffentlicht: (2024)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
von: Luo, Yuchen, et al.
Veröffentlicht: (2026)
von: Luo, Yuchen, et al.
Veröffentlicht: (2026)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-scale Approximate Nearest Neighbor Search
von: Tian, Bing, et al.
Veröffentlicht: (2024)
von: Tian, Bing, et al.
Veröffentlicht: (2024)
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
von: Feng, Laingjun, et al.
Veröffentlicht: (2025)
von: Feng, Laingjun, et al.
Veröffentlicht: (2025)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
von: Sui, Yang, et al.
Veröffentlicht: (2024)
von: Sui, Yang, et al.
Veröffentlicht: (2024)
BADA datasets
von: Hu, Xiabing
Veröffentlicht: (2026)
von: Hu, Xiabing
Veröffentlicht: (2026)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
von: Gao, Chang, et al.
Veröffentlicht: (2024)
von: Gao, Chang, et al.
Veröffentlicht: (2024)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
von: Prieto, Pablo, et al.
Veröffentlicht: (2025)
von: Prieto, Pablo, et al.
Veröffentlicht: (2025)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity
von: Tathe, Shivnath
Veröffentlicht: (2026)
von: Tathe, Shivnath
Veröffentlicht: (2026)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
Towards Efficient Multi-Scale Deformable Attention on NPU
von: Huang, Chenghuan, et al.
Veröffentlicht: (2025)
von: Huang, Chenghuan, et al.
Veröffentlicht: (2025)
Verification of Bit-Flip Attacks against Quantized Neural Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Q-BERT4Rec: Quantized Semantic-ID Representation Learning for Multimodal Recommendation
von: Huang, Haofeng, et al.
Veröffentlicht: (2025)
von: Huang, Haofeng, et al.
Veröffentlicht: (2025)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
von: Zhang, Jinhao Zhang Yunquan, et al.
Veröffentlicht: (2026)
von: Zhang, Jinhao Zhang Yunquan, et al.
Veröffentlicht: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
Bits and q-bits as versatility measures
von: José R.C. Piqueira
Veröffentlicht: (2004)
von: José R.C. Piqueira
Veröffentlicht: (2004)
BitVMX: A CPU for Universal Computation on Bitcoin
von: Lerner, Sergio Demian, et al.
Veröffentlicht: (2024)
von: Lerner, Sergio Demian, et al.
Veröffentlicht: (2024)
OSM+: Billion-Level OpenStreetMap Dataset for City-wide Experiments
von: Zheng, Guanjie, et al.
Veröffentlicht: (2025)
von: Zheng, Guanjie, et al.
Veröffentlicht: (2025)
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment
von: Sajid, Md. Ashiq Ul Islam, et al.
Veröffentlicht: (2026)
von: Sajid, Md. Ashiq Ul Islam, et al.
Veröffentlicht: (2026)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
Mixed-Precision Quantization: Make the Best Use of Bits Where They Matter Most
von: Fang, Yiming, et al.
Veröffentlicht: (2024)
von: Fang, Yiming, et al.
Veröffentlicht: (2024)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
von: Sharify, Sayeh, et al.
Veröffentlicht: (2026)
von: Sharify, Sayeh, et al.
Veröffentlicht: (2026)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
von: Gao, Jianyang, et al.
Veröffentlicht: (2024) -
Revisiting RaBitQ and TurboQuant: A Symmetric Comparison of Methods, Theory, and Experiments
von: Gao, Jianyang, et al.
Veröffentlicht: (2026) -
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
von: Shi, Jifan, et al.
Veröffentlicht: (2026) -
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
von: Wu, Jiehao, et al.
Veröffentlicht: (2026) -
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2026)