Matmul or No Matmul in the Era of 1-bit LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Malekar, Jinendra, Elbtity, Mohammed E., Zand, Ramtin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMPi: Optimizing LLMs for High-Throughput on Raspberry Pi
by: Ardakani, Mahsa, et al.
Published: (2025)
by: Ardakani, Mahsa, et al.
Published: (2025)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
by: Elbtity, Mohammed, et al.
Published: (2024)
by: Elbtity, Mohammed, et al.
Published: (2024)
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
by: Thangamani, Arun, et al.
Published: (2025)
by: Thangamani, Arun, et al.
Published: (2025)
Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural Networks
by: King, Grant, et al.
Published: (2025)
by: King, Grant, et al.
Published: (2025)
NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly Pipelines
by: Shyalika, Chathurangi, et al.
Published: (2025)
by: Shyalika, Chathurangi, et al.
Published: (2025)
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
by: Elhoushi, Mostafa, et al.
Published: (2025)
by: Elhoushi, Mostafa, et al.
Published: (2025)
Inverse Reinforcement Learning by Estimating Expertise of Demonstrators
by: Beliaev, Mark, et al.
Published: (2024)
by: Beliaev, Mark, et al.
Published: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024)
by: Mohammadi, Seyedali, et al.
Published: (2024)
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024)
by: Seekings, James, et al.
Published: (2024)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
by: Nielsen, Jacob, et al.
Published: (2025)
by: Nielsen, Jacob, et al.
Published: (2025)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
by: Lee, Jung Hyun, et al.
Published: (2025)
by: Lee, Jung Hyun, et al.
Published: (2025)
Conflict-Aware Adversarial Training
by: Xue, Zhiyu, et al.
Published: (2024)
by: Xue, Zhiyu, et al.
Published: (2024)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
by: Xia, Junhao, et al.
Published: (2025)
by: Xia, Junhao, et al.
Published: (2025)
4bit-Quantization in Vector-Embedding for RAG
by: Jeong, Taehee
Published: (2025)
by: Jeong, Taehee
Published: (2025)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Towards Low-bit Communication for Tensor Parallel LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
ICQuant: Index Coding enables Low-bit LLM Quantization
by: Li, Xinlin, et al.
Published: (2025)
by: Li, Xinlin, et al.
Published: (2025)
1.58-bit FLUX
by: Yang, Chenglin, et al.
Published: (2024)
by: Yang, Chenglin, et al.
Published: (2024)
Low-bit Model Quantization for Deep Neural Networks: A Survey
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
Wikipedia in the Era of LLMs: Evolution and Risks
by: Huang, Siming, et al.
Published: (2025)
by: Huang, Siming, et al.
Published: (2025)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
by: Ji, Xiaodong, et al.
Published: (2025)
by: Ji, Xiaodong, et al.
Published: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models
by: Thapa, Jikesh, et al.
Published: (2025)
by: Thapa, Jikesh, et al.
Published: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
by: Chen, Mengzhao, et al.
Published: (2025)
by: Chen, Mengzhao, et al.
Published: (2025)
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
by: Zhang, Michael S., et al.
Published: (2025)
by: Zhang, Michael S., et al.
Published: (2025)
GPU-accelerated simulated annealing based on p-bits with real-world device-variability modeling
by: Onizawa, Naoya, et al.
Published: (2026)
by: Onizawa, Naoya, et al.
Published: (2026)
Marconi: Prefix Caching for the Era of Hybrid LLMs
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows
by: Babaeipour, Ramtin, et al.
Published: (2026)
by: Babaeipour, Ramtin, et al.
Published: (2026)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
by: Yun, Juyoung, et al.
Published: (2023)
by: Yun, Juyoung, et al.
Published: (2023)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)
by: Xia, Haojun, et al.
Published: (2025)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
by: Han, Ruichi, et al.
Published: (2026)
by: Han, Ruichi, et al.
Published: (2026)
Rethinking Explainability in the Era of Multimodal AI
by: Agarwal, Chirag
Published: (2025)
by: Agarwal, Chirag
Published: (2025)
1 bit is all we need: binary normalized neural networks
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
Similar Items
-
LLMPi: Optimizing LLMs for High-Throughput on Raspberry Pi
by: Ardakani, Mahsa, et al.
Published: (2025) -
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025) -
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
by: Elbtity, Mohammed, et al.
Published: (2024) -
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
by: Thangamani, Arun, et al.
Published: (2025) -
Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural Networks
by: King, Grant, et al.
Published: (2025)