Efficient Hybrid Amplitude-Phase Quantization for Multi-Antenna Relay System
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Changdae, Jin, Xianglan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
von: Wang, Yulei, et al.
Veröffentlicht: (2025)
von: Wang, Yulei, et al.
Veröffentlicht: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
On General Linearly Implicit Quantized State System Methods
von: Bergonzi, Mariana, et al.
Veröffentlicht: (2025)
von: Bergonzi, Mariana, et al.
Veröffentlicht: (2025)
Energy-Efficient Software Development: A Multi-dimensional Empirical Analysis of Stack Overflow
von: Jin, Bihui, et al.
Veröffentlicht: (2024)
von: Jin, Bihui, et al.
Veröffentlicht: (2024)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
Beamforming-based Achievable Rate Maximization in ISAC System for Multi-UAV Networking
von: Zhou, Shengcai, et al.
Veröffentlicht: (2025)
von: Zhou, Shengcai, et al.
Veröffentlicht: (2025)
A Novel Hybrid Optical and STAR IRS System for NTN Communications
von: Shang, Shunyuan, et al.
Veröffentlicht: (2025)
von: Shang, Shunyuan, et al.
Veröffentlicht: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
AI Work Quantization Model: Closed-System AI Computational Effort Metric
von: Sharma, Aasish Kumar, et al.
Veröffentlicht: (2025)
von: Sharma, Aasish Kumar, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
von: Chai, Yuji, et al.
Veröffentlicht: (2025)
von: Chai, Yuji, et al.
Veröffentlicht: (2025)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
von: Jali, Neharika, et al.
Veröffentlicht: (2024)
von: Jali, Neharika, et al.
Veröffentlicht: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
Scaler: Efficient and Effective Cross Flow Analysis
von: Steven, et al.
Veröffentlicht: (2024)
von: Steven, et al.
Veröffentlicht: (2024)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
Efficient Data-Driven Production Scheduling in Pharmaceutical Manufacturing
von: Balatsos, Ioannis, et al.
Veröffentlicht: (2026)
von: Balatsos, Ioannis, et al.
Veröffentlicht: (2026)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
von: Wang, Tuowei, et al.
Veröffentlicht: (2026)
von: Wang, Tuowei, et al.
Veröffentlicht: (2026)
Reducing Waiting Time for Medical Tourists Through Hybrid Agent-Based and Discrete-Event Simulation: A Hospital Case Study
von: Baghi, Melika, et al.
Veröffentlicht: (2026)
von: Baghi, Melika, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
von: Zhao, Xinlong, et al.
Veröffentlicht: (2025)
von: Zhao, Xinlong, et al.
Veröffentlicht: (2025)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
von: Liu, Youyuan, et al.
Veröffentlicht: (2024)
von: Liu, Youyuan, et al.
Veröffentlicht: (2024)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
von: Li, Jianhui, et al.
Veröffentlicht: (2023)
von: Li, Jianhui, et al.
Veröffentlicht: (2023)
Multi-Strided Access Patterns to Boost Hardware Prefetching
von: Blom, Miguel O., et al.
Veröffentlicht: (2024)
von: Blom, Miguel O., et al.
Veröffentlicht: (2024)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
von: Sun, Jiaqi, et al.
Veröffentlicht: (2025)
von: Sun, Jiaqi, et al.
Veröffentlicht: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Attributing the System's Overall Effect to its Components
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
von: Ray, Kaustabha, et al.
Veröffentlicht: (2025)
von: Ray, Kaustabha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
von: Wang, Yulei, et al.
Veröffentlicht: (2025) -
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025) -
On General Linearly Implicit Quantized State System Methods
von: Bergonzi, Mariana, et al.
Veröffentlicht: (2025) -
Energy-Efficient Software Development: A Multi-dimensional Empirical Analysis of Stack Overflow
von: Jin, Bihui, et al.
Veröffentlicht: (2024) -
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)