ITQ3_S: High-Fidelity 3-bit LLM Inference via Interleaved Ternary Quantization with Rotation-Domain Smoothing
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Yoon, Edward J. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
par: Li, Zonghang, et autres
Publié: (2025)
par: Li, Zonghang, et autres
Publié: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
par: Zhang, Lingzhe, et autres
Publié: (2025)
par: Zhang, Lingzhe, et autres
Publié: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
par: Mohammad, Noor Islam S.
Publié: (2025)
par: Mohammad, Noor Islam S.
Publié: (2025)
A new Dune grid for scalable dynamic adaptivity based on the p4est software library
par: Burstedde, Carsten, et autres
Publié: (2025)
par: Burstedde, Carsten, et autres
Publié: (2025)
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
par: Li, Zonghang, et autres
Publié: (2024)
par: Li, Zonghang, et autres
Publié: (2024)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
Distributed Tomographic Reconstruction with Quantization
par: Miao, Runxuan, et autres
Publié: (2024)
par: Miao, Runxuan, et autres
Publié: (2024)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
par: Krishna, Gutha Jaya
Publié: (2024)
par: Krishna, Gutha Jaya
Publié: (2024)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
par: Sada, Mohammad Firas, et autres
Publié: (2025)
par: Sada, Mohammad Firas, et autres
Publié: (2025)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
par: Taubert, Oskar, et autres
Publié: (2023)
par: Taubert, Oskar, et autres
Publié: (2023)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
par: Suzuki, Kengo, et autres
Publié: (2026)
par: Suzuki, Kengo, et autres
Publié: (2026)
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
par: Fernandes, Daniel, et autres
Publié: (2025)
par: Fernandes, Daniel, et autres
Publié: (2025)
Parallel Quadratic Selected Inversion in Quantum Transport Simulation
par: Maillou, Vincent, et autres
Publié: (2026)
par: Maillou, Vincent, et autres
Publié: (2026)
Efficient Parallel Scheduling for Sparse Triangular Solvers
par: Böhnlein, Toni, et autres
Publié: (2025)
par: Böhnlein, Toni, et autres
Publié: (2025)
LLM-Viterbi: Semantic-Aware Decoding for Convolutional Codes
par: Li, Zhengtong, et autres
Publié: (2026)
par: Li, Zhengtong, et autres
Publié: (2026)
Synthesis of signal processing algorithms with constraints on minimal parallelism and memory space
par: Salishev, Sergey
Publié: (2025)
par: Salishev, Sergey
Publié: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
par: Gao, Jianhua, et autres
Publié: (2024)
par: Gao, Jianhua, et autres
Publié: (2024)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
par: Dang, Kieu, et autres
Publié: (2025)
par: Dang, Kieu, et autres
Publié: (2025)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
par: Abzianidze, Lasha
Publié: (2023)
par: Abzianidze, Lasha
Publié: (2023)
On Reduction and Synthesis of Petri's Cycloids
par: Valk, Rüdiger, et autres
Publié: (2025)
par: Valk, Rüdiger, et autres
Publié: (2025)
Modelling cooperating failure-resilient Processes
par: Valk, Rüdiger
Publié: (2024)
par: Valk, Rüdiger
Publié: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
par: Anam, Rizal Khoirul
Publié: (2025)
par: Anam, Rizal Khoirul
Publié: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
par: Park, Seungcheol, et autres
Publié: (2025)
par: Park, Seungcheol, et autres
Publié: (2025)
A multigrid reduction framework for domains with symmetries
par: Alsalti-Baldellou, Àdel, et autres
Publié: (2024)
par: Alsalti-Baldellou, Àdel, et autres
Publié: (2024)
The Performance of Low-Synchronization Variants of Reorthogonalized Block Classical Gram--Schmidt
par: Carson, Erin, et autres
Publié: (2025)
par: Carson, Erin, et autres
Publié: (2025)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
par: Gao, Jianhua, et autres
Publié: (2024)
par: Gao, Jianhua, et autres
Publié: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
par: Gao, Jianhua, et autres
Publié: (2024)
par: Gao, Jianhua, et autres
Publié: (2024)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
par: Otal, Hakan T., et autres
Publié: (2024)
par: Otal, Hakan T., et autres
Publié: (2024)
Math Natural Language Inference: this should be easy!
par: de Paiva, Valeria, et autres
Publié: (2025)
par: de Paiva, Valeria, et autres
Publié: (2025)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
par: Schroeder, Christian, et autres
Publié: (2023)
par: Schroeder, Christian, et autres
Publié: (2023)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
par: Liu, Aiwei, et autres
Publié: (2025)
par: Liu, Aiwei, et autres
Publié: (2025)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
par: Zanbaghi, Shahin, et autres
Publié: (2025)
par: Zanbaghi, Shahin, et autres
Publié: (2025)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
par: Pan, Leyi, et autres
Publié: (2025)
par: Pan, Leyi, et autres
Publié: (2025)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
par: Muñoz-Ortiz, Alberto, et autres
Publié: (2023)
par: Muñoz-Ortiz, Alberto, et autres
Publié: (2023)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
par: Johnson, Warren
Publié: (2026)
par: Johnson, Warren
Publié: (2026)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
par: Wang, Yongjie, et autres
Publié: (2025)
par: Wang, Yongjie, et autres
Publié: (2025)
Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth
par: Nair, Arjun S.
Publié: (2026)
par: Nair, Arjun S.
Publié: (2026)
Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
par: Ethiraj, Vignesh, et autres
Publié: (2025)
par: Ethiraj, Vignesh, et autres
Publié: (2025)
Documents similaires
-
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
par: Li, Zonghang, et autres
Publié: (2025) -
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
par: Zhang, Lingzhe, et autres
Publié: (2025) -
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
par: Mohammad, Noor Islam S.
Publié: (2025) -
A new Dune grid for scalable dynamic adaptivity based on the p4est software library
par: Burstedde, Carsten, et autres
Publié: (2025) -
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
par: Li, Zonghang, et autres
Publié: (2024)